Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Oct 22, 2024

Xiaoqian Shen, Yunyang Xiong, Changsheng Zhao, Lemeng Wu, Jun Chen, Chenchen Zhu, Zechun Liu, Fanyi Xiao, Balakrishnan Varadarajan, Florian Bordes(+7 more)

Figure 1 for LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Figure 2 for LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Figure 3 for LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Figure 4 for LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Share this with someone who'll enjoy it:

Abstract:Multimodal Large Language Models (MLLMs) have shown promising progress in understanding and analyzing video content. However, processing long videos remains a significant challenge constrained by LLM's context size. To address this limitation, we propose LongVU, a spatiotemporal adaptive compression mechanism thats reduces the number of video tokens while preserving visual details of long videos. Our idea is based on leveraging cross-modal query and inter-frame dependencies to adaptively reduce temporal and spatial redundancy in videos. Specifically, we leverage DINOv2 features to remove redundant frames that exhibit high similarity. Then we utilize text-guided cross-modal query for selective frame feature reduction. Further, we perform spatial token reduction across frames based on their temporal dependencies. Our adaptive compression strategy effectively processes a large number of frames with little visual information loss within given context length. Our LongVU consistently surpass existing methods across a variety of video understanding benchmarks, especially on hour-long video understanding tasks such as VideoMME and MLVU. Given a light-weight LLM, our LongVU also scales effectively into a smaller size with state-of-the-art video understanding performance.

* Project page: https://vision-cair.github.io/LongVU

View paper on

Share this with someone who'll enjoy it:

Title:LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Paper and Code