Conditional Image-to-Video Generation with Latent Flow Diffusion Models
Haomiao Ni, Changhao Shi, Kai Li, Sharon X. Huang, Martin Renqiang Min
摘要
Conditional image-to-video (cI2V) generation aims to synthesize a new plausible video starting from an image (e.g., a person's face) and a condition (e.g., an action class label like smile). The key challenge of the cI2V task lies in the simultaneous generation of realistic spatial appearance and temporal dynamics corresponding to the given image and condition. In this paper, we propose an approach for cI2V using novel latent flow diffusion models (LFDM) that synthesize an optical flow sequence in the latent space based on the given condition to warp the given image. Compared to previous direct-synthesis-based works, our proposed LFDM can better synthesize spatial details and temporal motion by fully utilizing the spatial content of the given image and warping it in the latent space according to the generated temporally-coherent flow. The training of LFDM consists of two separate stages: (1) an unsupervised learning stage to train a latent flow auto-encoder for spatial content generation, including a flow predictor to estimate latent flow between pairs of video frames, and (2) a conditional learning stage to train a 3D-UNet-based diffusion model (DM) for temporal latent flow generation. Unlike previous DMs operating in pixel space or latent feature space that couples spatial and temporal information, the DM in our LFDM only needs to learn a low-dimensional latent flow space for motion generation, thus being more computationally efficient. We conduct comprehensive experiments on multiple datasets, where LFDM consistently outperforms prior arts. Furthermore, we show that LFDM can be easily adapted to new domains by simply finetuning the image decoder. Our code is available at https: //github.com/nihaomiao/CVPR23_LFDM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper76
- VideoComposer: Compositional Video Synthesis with Motion ControllabilityXiang Wang, Hangjie Yuan, Shiwei Zhang, Dayou Chen 等NeurIPS 2023 · 被引用 579 次
- StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationYupeng Zhou, Daquan Zhou, Ming-Ming Cheng, Jiashi Feng 等NeurIPS 2024 · 被引用 291 次
- PreDiff: Precipitation Nowcasting with Latent Diffusion ModelsZhihan Gao, Xingjian Shi, Boran Han, Hao Wang 等NeurIPS 2023 · 被引用 171 次
- Disco: Disentangled Control for Realistic Human Dance GenerationTan Wang, Linjie Li, Kevin Lin, Yuanhao Zhai 等CVPR 2024 · 被引用 62 次
- MotionCraft: Physics-Based Zero-Shot Video GenerationAntonio Montanaro, Luca Savant Aira, Emanuele Aiello, Diego Valsesia 等NeurIPS 2024 · 被引用 52 次
它引用的顶会 Paper29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- Decouple Content and Motion for Conditional Image-to-Video GenerationCuifeng Shen, Yulu Gan, Chen Chen, Xiongwei Zhu 等AAAI 2024 · 被引用 13 次
- Efficient Video Diffusion Models via Content-Frame Motion-Latent DecompositionSihyun Yu, Weili Nie, De-An Huang, Boyi Li 等ICLR 2024 · 被引用 34 次
- Realtime Video Frame Interpolation using One-Step Diffusion SamplingYongrui Ma, Shijie Zhao, Mingde Yao, Junlin Li 等ICLR 2026
- Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion ModelMin Zhao, Hongzhou Zhu, Chendong Xiang, Kaiwen Zheng 等NeurIPS 2024 · 被引用 33 次
- Video Probabilistic Diffusion Models in Projected Latent SpaceSihyun Yu, Kihyuk Sohn, Subin Kim, Jinwoo ShinCVPR 2023
