VideoGuide: Improving Video Diffusion Models without Training Through a Teacher's Guide
Dohun Lee, Bryan Sangwoo Kim, Geon Yeong Park, Jong Chul Ye
Abstract
Text-to-image (T2I) diffusion models have revolutionized visual content creation, but extending these capabilities to textto-video (T2V) generation remains a challenge, particularly in preserving temporal consistency. Existing methods that aim to improve consistency often cause trade-offs such as reduced imaging quality and impractical computational time.
To address these issues we introduce VideoGuide, a novel framework that enhances the temporal consistency of pretrained T2V models without the need for additional training or fine-tuning. Instead, VideoGuide leverages any pretrained video diffusion model (VDM) or itself as a guide during the early stages of inference, improving temporal quality by interpolating the guiding model's denoised samples into the sampling model's denoising process. The proposed method
This CVPR paper is the Open Access version, provided by the Computer Vision Foundation. Except for this watermark, it is identical to the accepted version; the final published version of the proceedings is available on IEEE Xplore.
brings about significant improvement in temporal consistency and image fidelity, providing a cost-effective and practical solution that synergizes the strengths of various video diffusion models. Furthermore, we demonstrate prior distillation, revealing that base models can achieve enhanced text coherence by utilizing the superior data prior of the guiding model through the proposed method. Project Page: https: //dohunlee1.github.io/videoguide.github.io/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acf755fd-d209-4478-9932-c831bcf6f4b5Cited by top-tier papers4
- FlowMo: Variance-Based Flow Guidance for Coherent Motion in Video GenerationAriel Shaulov, Itay Hazan, Lior Wolf, Hila CheferNeurIPS 2025 · 22 citations
- Motion Prior Distillation in Time Reversal Sampling for Generative InbetweeningWooseok Jeon, Seunghyun Shin, Dongmin Shin, Hae-Gon JeonICLR 2026 · 5 citations
- Optical-Flow Guided Prompt Optimization for Coherent Video GenerationHyelin Nam, Jaemin Kim, Dohun Lee, Jong Chul YeCVPR 2025
- Trajectory-Stabilized Inference for Diffusion-Based Video InpaintingZhanhe Zhang, Jiahua Li, Xu Yang, Kun Wei et al.ICML 2026
Builds on24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- T2V-Turbo-v2: Enhancing Video Model Post-Training through Data, Reward, and Conditional Guidance DesignJiachen Li, Qian Long, Jian Zheng, Xiaofeng Gao et al.ICLR 2025
- COVE: Unleashing the Diffusion Feature Correspondence for Consistent Video EditingJiangshan Wang, Yue Ma, Jiayi Guo, Yicheng Xiao et al.NeurIPS 2024 · 76 citations
- MeDM: Mediating Image Diffusion Models for Video-to-Video Translation with Temporal Correspondence GuidanceErnie Chu, Tzuhsuan Huang, Shuo-Yen Lin, Jun-Cheng ChenAAAI 2024 · 25 citations
- TokenFlow: Consistent Diffusion Features for Consistent Video EditingMichal Geyer, Omer Bar-Tal, Shai Bagon, Tali DekelICLR 2024 · 439 citations
- Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-ResolutionShangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo et al.CVPR 2024 · 52 citations
