Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Oct 03, 2024

Yuqing Wang, Tianwei Xiong, Daquan Zhou, Zhijie Lin, Yang Zhao, Bingyi Kang, Jiashi Feng, Xihui Liu

Figure 1 for Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Figure 2 for Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Figure 3 for Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Figure 4 for Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Share this with someone who'll enjoy it:

Abstract:It is desirable but challenging to generate content-rich long videos in the scale of minutes. Autoregressive large language models (LLMs) have achieved great success in generating coherent and long sequences of tokens in the domain of natural language processing, while the exploration of autoregressive LLMs for video generation is limited to generating short videos of several seconds. In this work, we conduct a deep analysis of the challenges that prevent autoregressive LLM-based video generators from generating long videos. Based on the observations and analysis, we propose Loong, a new autoregressive LLM-based video generator that can generate minute-long videos. Specifically, we model the text tokens and video tokens as a unified sequence for autoregressive LLMs and train the model from scratch. We propose progressive short-to-long training with a loss re-weighting scheme to mitigate the loss imbalance problem for long video training. We further investigate inference strategies, including video token re-encoding and sampling strategies, to diminish error accumulation during inference. Our proposed Loong can be trained on 10-second videos and be extended to generate minute-level long videos conditioned on text prompts, as demonstrated by the results. More samples are available at: https://epiphqny.github.io/Loong-video.

* Project page: https://epiphqny.github.io/Loong-video/

View paper on

Share this with someone who'll enjoy it:

Title:Loong: Generating Minute-level Long Videos with Autoregressive Language Models

Paper and Code