Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Decouple Content and Motion for Conditional Image-to-Video Generation

Nov 24, 2023

Cuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu, Lele Cheng, Jinzhi Wang

Figure 1 for Decouple Content and Motion for Conditional Image-to-Video Generation

Figure 2 for Decouple Content and Motion for Conditional Image-to-Video Generation

Figure 3 for Decouple Content and Motion for Conditional Image-to-Video Generation

Figure 4 for Decouple Content and Motion for Conditional Image-to-Video Generation

Share this with someone who'll enjoy it:

Abstract:The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text.The previous cI2V generation methods conventionally perform in RGB pixel space, with limitations in modeling motion consistency and visual continuity. Additionally, the efficiency of generating videos in pixel space is quite low.In this paper, we propose a novel approach to address these challenges by disentangling the target RGB pixels into two distinct components: spatial content and temporal motions. Specifically, we predict temporal motions which include motion vector and residual based on a 3D-UNet diffusion model. By explicitly modeling temporal motions and warping them to the starting image, we improve the temporal consistency of generated videos. This results in a reduction of spatial redundancy, emphasizing temporal details. Our proposed method achieves performance improvements by disentangling content and motion, all without introducing new structural complexities to the model. Extensive experiments on various datasets confirm our approach's superior performance over the majority of state-of-the-art methods in both effectiveness and efficiency.

View paper on

Share this with someone who'll enjoy it:

Title:Decouple Content and Motion for Conditional Image-to-Video Generation

Paper and Code