Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Zero-Shot Long-Form Video Understanding through Screenplay

Jun 25, 2024

Yongliang Wu, Bozheng Li, Jiawang Cao, Wenbo Zhu, Yi Lu, Weiheng Chi, Chuyun Xie, Haolin Zheng, Ziyue Su, Jay Wu(+1 more)

Figure 1 for Zero-Shot Long-Form Video Understanding through Screenplay

Figure 2 for Zero-Shot Long-Form Video Understanding through Screenplay

Figure 3 for Zero-Shot Long-Form Video Understanding through Screenplay

Figure 4 for Zero-Shot Long-Form Video Understanding through Screenplay

Share this with someone who'll enjoy it:

Abstract:The Long-form Video Question-Answering task requires the comprehension and analysis of extended video content to respond accurately to questions by utilizing both temporal and contextual information. In this paper, we present MM-Screenplayer, an advanced video understanding system with multi-modal perception capabilities that can convert any video into textual screenplay representations. Unlike previous storytelling methods, we organize video content into scenes as the basic unit, rather than just visually continuous shots. Additionally, we developed a ``Look Back'' strategy to reassess and validate uncertain information, particularly targeting breakpoint mode. MM-Screenplayer achieved highest score in the CVPR'2024 LOng-form VidEo Understanding (LOVEU) Track 1 Challenge, with a global accuracy of 87.5% and a breakpoint accuracy of 68.8%.

* Highest Score Award to the CVPR'2024 LOVEU Track 1 Challenge

View paper on

Share this with someone who'll enjoy it:

Title:Zero-Shot Long-Form Video Understanding through Screenplay

Paper and Code