ControlVideo: Training-free Controllable Text-to-video Generation
Yabo Zhang, Yuxiang Wei, Dongsheng Jiang, Xiaopeng Zhang, Wangmeng Zuo, Qi Tian
摘要
Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temporal modeling. Besides the training burden, the generated videos also suffer from appearance inconsistency and structural flickers, especially in long video synthesis. To address these challenges, we design a training-free framework called ControlVideo to enable natural and efficient text-to-video generation. ControlVideo, adapted from ControlNet, leverages coarsely structural consistency from input motion sequences, and introduces three modules to improve video generation. Firstly, to ensure appearance coherence between frames, ControlVideo adds fully cross-frame interaction in self-attention modules. Secondly, to mitigate the flicker effect, it introduces an interleaved-frame smoother that employs frame interpolation on alternated frames. Finally, to produce long videos efficiently, it utilizes a hierarchical sampler that separately synthesizes each short clip with holistic coherency. Empowered with these modules, ControlVideo outperforms the state-of-the-arts on extensive motion-prompt pairs quantitatively and qualitatively. Notably, thanks to the efficient designs, it generates both short and long videos within several minutes using one NVIDIA 2080Ti. Code is available at https://github.com/YBYBZhang/ControlVideo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper139
- iVideoGPT: Interactive VideoGPTs are Scalable World ModelsJialong Wu, Shaofeng Yin, Ningya Feng, Xu He 等NeurIPS 2024 · 被引用 177 次
- FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editingYuren Cong, Mengmeng Xu, Christian Simon, Shoufa Chen 等ICLR 2024 · 被引用 175 次
- Frequency Perception Network for Camouflaged Object DetectionRunmin Cong, Mengyao Sun, Sanyi Zhang, Xiaofei Zhou 等ACM MM 2023 · 被引用 130 次
- ReVideo: Remake a Video with Motion and Content ControlChong Mou, Mingdeng Cao, Xintao Wang, Zhaoyang Zhang 等NeurIPS 2024 · 被引用 87 次
- Dual Data Alignment Makes AI-Generated Image Detector Easier GeneralizableRuoxin Chen, Junwei Xi, Zhiyuan Yan, Ke-Yue Zhang 等NeurIPS 2025 · 被引用 78 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
相关 Paper
- ConditionVideo: Training-Free Condition-Guided Video GenerationBo Peng, Xinyuan Chen, Yaohui Wang, Chaochao Lu 等AAAI 2024 · 被引用 7 次
- Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion ModelHan Lin, Jaemin Cho, Abhay Zala, Mohit BansalICLR 2025
- Fuse Your Latents: Video Editing with Multi-source Latent Diffusion ModelsTianyi Lu, Xing Zhang, Jiaxi Gu, Renjing Pei 等ACM MM 2024 · 被引用 2 次
- MotionFlow: Attention-Driven Motion Transfer in Video Diffusion ModelsTuna Han Salih Meral, Hidir Yesiltepe, Connor Dunlop, Pinar YanardagAAAI 2026
- LongDiff: Training-Free Long Video Generation in One GoZhuoling Li, Hossein Rahmani, Qiuhong Ke, Jun LiuCVPR 2025
