Dual-Expert Consistency Model for Efficient and High-Quality Video Generation
Zhengyao Lv, Chenyang Si, Tianlin Pan, Zhaoxi Chen, Kwan-Yee K. Wong, Yu Qiao, Ziwei Liu
摘要
Diffusion Models have achieved remarkable results in video synthesis but require iterative denoising steps, leading to substantial computational overhead. Consistency Models have made significant progress in accelerating diffusion models. However, directly applying them to video diffusion models often results in severe degradation of temporal consistency and appearance details. In this paper, by analyzing the training dynamics of Consistency Models, we identify a key conflicting learning dynamics during the distillation process: there is a significant discrepancy in the optimization gradients and loss contributions across different timesteps. This discrepancy prevents the distilled student model from achieving an optimal state, leading to compromised temporal consistency and degraded appearance details. To address this issue, we propose a parameter-efficient Dual-Expert Consistency Model (DCM), where a semantic expert focuses on learning semantic layout and motion, while a detail expert specializes in fine detail refinement. Furthermore, we introduce Temporal Coherence Loss to improve motion consistency for the semantic expert and apply GAN and Feature Matching Loss to enhance the synthesis quality of the detail expert. Our approach achieves state-of-the-art visual quality with significantly reduced sampling steps, demonstrating the effectiveness of expert specialization in video diffusion model distillation. Our code and models are available at https://github.com/Vchitect/DCM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step GenerationYuyang You, Yongzhi Li, Jiahui Li, Yadong Mu 等CVPR 2026 · 被引用 7 次
- FlashMotion: Few-Step Controllable Video Generation with Trajectory GuidanceQuanhao Li, Zhen Xing, Rui Wang, Haidong Cao 等CVPR 2026 · 被引用 5 次
- DUO-VSR: Dual-Stream Distillation for One-Step Video Super-ResolutionZhengyao Lv, Menghan Xia, Xintao Wang, Kwan-Yee K. WongCVPR 2026 · 被引用 4 次
- SwitchCraft: Training-Free Multi-Event Video Generation with Attention ControlsQianxun Xu, Chenxi Song, Yujun Cai, Chi ZhangCVPR 2026 · 被引用 4 次
- FlowSteer: Guiding Few-Step Image Synthesis with Authentic TrajectoriesLei Ke, Hubery Yin, Gongye Liu, Zhengyao Lv 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper49
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Large Scale Diffusion Distillation via Score-Regularized Continuous-Time ConsistencyKaiwen Zheng, Yuji Wang, Qianli Ma, Huayu Chen 等ICLR 2026 · 被引用 77 次
- Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance DistillationYuanhao Zhai, Kevin Lin, Zhengyuan Yang, Linjie Li 等NeurIPS 2024 · 被引用 41 次
- SwiftVideo: A Unified Framework for Few-Step Video Generation Through Trajectory-Distribution AlignmentYanxiao Sun, Jiafu Wu, Yun Cao, Chengming Xu 等AAAI 2026 · 被引用 6 次
- T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward FeedbackJiachen Li, Weixi Feng, Tsu-Jui Fu, Xinyi Wang 等NeurIPS 2024 · 被引用 97 次
- See Further When Clear: Curriculum Consistency ModelYunpeng Liu, Boxiao Liu, Yi Zhang, Xingzhong Hou 等CVPR 2025
