VideoElevator: Elevating Video Generation Quality with Versatile Text-to-Image Diffusion Models
Yabo Zhang, Yuxiang Wei, Xianhui Lin, Zheng Hui, Peiran Ren, Xuansong Xie, Wangmeng Zuo
Abstract
Text-to-image diffusion models (T2I) have demonstrated unprecedented capabilities in creating realistic and aesthetic images. On the contrary, text-to-video diffusion models (T2V) still lag far behind in frame quality and text alignment, owing to insufficient quality and quantity of training videos. In this paper, we introduce VideoElevator, a training-free and plug-and-play method, which elevates the performance of T2V using superior capabilities of T2I. Different from conventional T2V sampling (i.e., temporal and spatial modeling), VideoElevator explicitly decomposes each sampling step into temporal motion refining and spatial quality elevating. Specifically, temporal motion refining uses encapsulated T2V to enhance temporal consistency, followed by inverting to the noise distribution required by T2I. Then, spatial quality elevating harnesses inflated T2I to directly predict less noisy latent, adding more photo-realistic details. We have conducted experiments in extensive prompts under the combination of various T2V and T2I. The results show that VideoElevator not only improves the performance of T2V baselines with foundational T2I, but also facilitates stylistic video synthesis with personalized T2I. Please watch all videos in supplementary materials for better view.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 21e99b37-be78-47b8-a4c6-c2e82f000098Cited by top-tier papers7
- RefDrop: Controllable Consistency in Image or Video Generation via Reference Feature GuidanceJiaojiao Fan, Haotian Xue, Qinsheng Zhang, Yongxin ChenNeurIPS 2024 · 7 citations
- FramePainter: Endowing Interactive Image Editing with Video Diffusion PriorsYabo Zhang, Xinpeng Zhou, Yihan Zeng, Hang Xu et al.ICCV 2025 · 3 citations
- Diffusion-Driven Progressive Target Manipulation for Source-Free Domain AdaptationYuyang Huang, Yabo Chen, Junyu Zhou, Wenrui Dai et al.NeurIPS 2025 · 2 citations
- Dual-Expert Consistency Model for Efficient and High-Quality Video GenerationZhengyao Lv, Chenyang Si, Tianlin Pan, Zhaoxi Chen et al.ICCV 2025 · 1 citation
- Encapsulated Composition of Text-to-Image and Text-to-Video Models for High-Quality Video SynthesisTongtong Su, Chengyu Wang, Bingyan Liu, Jun Huang et al.CVPR 2025
Builds on43
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Fuse Your Latents: Video Editing with Multi-source Latent Diffusion ModelsTianyi Lu, Xing Zhang, Jiaxi Gu, Renjing Pei et al.ACM MM 2024 · 2 citations
- Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You ThinkJie Tian, Xiaoye Qu, Zhenyi Lu, Wei Wei et al.CVPR 2025
- Accelerating Diffusion Sampling via Exploiting Local Transition CoherenceShangwen Zhu, Han Zhang, Zhantao Yang, Qianyu Peng et al.ICCV 2025
- Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock DenoisingAssaf Singer, Noam Rotstein, Amir Mann, Ron Kimmel et al.ICLR 2026 · 13 citations
- FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video GenerationAriel Shaulov, Itay Hazan, Lior Wolf, Hila CheferNeurIPS 2025 · 22 citations
