Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency
Kaiwen Zheng, Yuji Wang, Qianli Ma, Huayu Chen, Jintao Zhang, Yogesh Balaji, Jianfei Chen, Ming-Yu Liu, Jun Zhu, Qinsheng Zhang
Abstract
Although continuous-time consistency models (e.g., sCM, MeanFlow) are theoretically principled and empirically powerful for fast academic-scale diffusion, its applicability to large-scale text-to-image and video tasks remains unclear due to infrastructure challenges in Jacobian-vector product (JVP) computation and the limitations of evaluation benchmarks like FID. This work represents the first effort to scale up continuous-time consistency to general application-level image and video diffusion models, and to make JVP-based distillation effective at large scale. We first develop a parallelism-compatible FlashAttention-2 JVP kernel, enabling sCM training on models with over 10 billion parameters and high-dimensional video tasks. Our investigation reveals fundamental quality limitations of sCM in fine-detail generation, which we attribute to error accumulation and the “mode-covering” nature of its forward-divergence objective. To remedy this, we propose the score-regularized continuous-time consistency model (rCM), which incorporates score distillation as a long-skip regularizer. This integration complements sCM with the “mode-seeking” reverse divergence, effectively improving visual quality while maintaining high generation diversity. Validated on large-scale models (Cosmos-Predict2, Wan2.1) up to 14B parameters and 5-second videos, rCM generally matches the state-of-the-art distillation method DMD2 on quality metrics while mitigating mode collapse and offering notable advantages in diversity, all without GAN tuning or extensive hyperparameter searches. The distilled models generate high-fidelity samples in only steps, accelerating diffusion sampling by . These results position rCM as a practical and theoretically grounded framework for advancing large-scale diffusion distillation. Code is available at https://github.com/NVlabs/rcm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80cdec3a-525d-498e-bbdd-36d2147aca8bCited by top-tier papers20
- pi-Flow: Policy-Based Few-Step Generation via Imitation DistillationHansheng Chen, Kai Zhang, Hao Tan, Leonidas Guibas et al.ICLR 2026 · 26 citations
- Transition Matching Distillation for Fast Video GenerationWeili Nie, Julius Berner, Nanye Ma, Chao Liu et al.CVPR 2026 · 24 citations
- TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial FlowsZhenglin Cheng, Peng Sun, Jianguo Li, Tao LinICLR 2026 · 17 citations
- Flow Map Distillation Without DataShangyuan Tong, Nanye Ma, Saining Xie, Tommi S. JaakkolaCVPR 2026 · 13 citations
- Diversity-Preserved Distribution Matching Distillation for Fast Visual SynthesisTianhe Wu, Ruibin Li, Lei Zhang, Kede MaICML 2026 · 12 citations
Related papers
- LogCD: Local-to-global Consistency Distillation for Few-step Image GenerationQingsong Xie, Zhenyi Liao, Chen Chen, Zhijie Deng et al.CVPR 2026
- Dual-Expert Consistency Model for Efficient and High-Quality Video GenerationZhengyao Lv, Chenyang Si, Tianlin Pan, Zhaoxi Chen et al.ICCV 2025 · 1 citation
- Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement LearningGuanjie Chen, Shirui Huang, Yifu Sun, Kai Liu et al.CVPR 2026 · 17 citations
- T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward FeedbackJiachen Li, Weixi Feng, Tsu-Jui Fu, Xinyi Wang et al.NeurIPS 2024 · 97 citations
- OSV: One Step is Enough for High-Quality Image to Video GenerationXiaofeng Mao, Zhengkai Jiang, Fu-Yun Wang, Jiangning Zhang et al.CVPR 2025
