Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Tiehan Fan

InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

Dec 12, 2024

Tiehan Fan, Kepan Nan, Rui Xie, Penghao Zhou, Zhenheng Yang, Chaoyou Fu, Xiang Li, Jian Yang, Ying Tai

Figure 1 for InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

Figure 2 for InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

Figure 3 for InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

Figure 4 for InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption

Abstract:Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However, current video captions often suffer from insufficient details, hallucinations and imprecise motion depiction, affecting the fidelity and consistency of generated videos. In this work, we propose a novel instance-aware structured caption framework, termed InstanceCap, to achieve instance-level and fine-grained video caption for the first time. Based on this scheme, we design an auxiliary models cluster to convert original video into instances to enhance instance fidelity. Video instances are further used to refine dense prompts into structured phrases, achieving concise yet precise descriptions. Furthermore, a 22K InstanceVid dataset is curated for training, and an enhancement pipeline that tailored to InstanceCap structure is proposed for inference. Experimental results demonstrate that our proposed InstanceCap significantly outperform previous models, ensuring high fidelity between captions and videos while reducing hallucinations.

Via

Access Paper or Ask Questions

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Jul 02, 2024

Kepan Nan, Rui Xie, Penghao Zhou, Tiehan Fan, Zhenheng Yang, Zhijie Chen, Xiang Li, Jian Yang, Ying Tai

Figure 1 for OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Figure 2 for OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Figure 3 for OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Figure 4 for OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation

Abstract:Text-to-video (T2V) generation has recently garnered significant attention thanks to the large multi-modality model Sora. However, T2V generation still faces two important challenges: 1) Lacking a precise open sourced high-quality dataset. The previous popular video datasets, e.g. WebVid-10M and Panda-70M, are either with low quality or too large for most research institutions. Therefore, it is challenging but crucial to collect a precise high-quality text-video pairs for T2V generation. 2) Ignoring to fully utilize textual information. Recent T2V methods have focused on vision transformers, using a simple cross attention module for video generation, which falls short of thoroughly extracting semantic information from text prompt. To address these issues, we introduce OpenVid-1M, a precise high-quality dataset with expressive captions. This open-scenario dataset contains over 1 million text-video pairs, facilitating research on T2V generation. Furthermore, we curate 433K 1080p videos from OpenVid-1M to create OpenVidHD-0.4M, advancing high-definition video generation. Additionally, we propose a novel Multi-modal Video Diffusion Transformer (MVDiT) capable of mining both structure information from visual tokens and semantic information from text tokens. Extensive experiments and ablation studies verify the superiority of OpenVid-1M over previous datasets and the effectiveness of our MVDiT.

* 15 pages, 9 figures

Via

Access Paper or Ask Questions