DOLLAR: Few-Step Video Generation Via Distillation and Latent Reward Optimization
Zihan Ding, Chi Jin, Difan Liu, Haitian Zheng, Krishna Kumar Singh, Qiang Zhang, Yan Kang, Zhe Lin, Yuchen Liu
Abstract
Diffusion probabilistic models have shown significant progress in video generation; however, their computational efficiency is limited by the large number of sampling steps required. Reducing sampling steps often compromises video quality or generation diversity. In this work, we introduce a distillation method that combines variational score distillation and consistency distillation to achieve few-step video generation, maintaining both high quality and diversity. We also propose a latent reward model fine-tuning approach to further enhance video generation performance according to any specified reward metric. This approach reduces memory usage and does not require the reward to be differentiable. Our method demonstrates state-of-the-art performance in few-step generation for 10-second videos (128 frames at 12 FPS). The distilled student model achieves a score of 82.57 on VBench, surpassing the teacher model as well as baseline models Gen-3, T2V-Turbo [25], and Kling [24]. One-step distillation accelerates the teacher model's diffusion sampling by up to 278.6 times, enabling near real-time generation. Human evaluations further validate the superior performance of our 4-step student models compared to teacher model using 50-step DDIM sampling.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Learning World Models for Interactive Video GenerationTaiye Chen, Xun Hu, Zihan Ding, Chi JinNeurIPS 2025 · 37 citations
- FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video GenerationShilong Zhang, Wenbo Li, Shoufa Chen, Chongjian Ge et al.AAAI 2026 · 30 citations
- Transition Matching Distillation for Fast Video GenerationWeili Nie, Julius Berner, Nanye Ma, Chao Liu et al.CVPR 2026 · 24 citations
- FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video DiffusionAkide Liu, Zeyu Zhang, Zhexin Li, Xuehai Bai et al.NeurIPS 2025 · 19 citations
- Learning from Noisy Preferences: A Semi-Supervised Learning Approach to Direct Preference OptimizationXinxin Liu, Ming Li, Zonglin Lyu, Yuzhang Shang et al.ICLR 2026 · 5 citations
Builds on44
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- T2V-Turbo: Breaking the Quality Bottleneck of Video Consistency Model with Mixed Reward FeedbackJiachen Li, Weixi Feng, Tsu-Jui Fu, Xinyi Wang et al.NeurIPS 2024 · 97 citations
- SwiftVideo: A Unified Framework for Few-Step Video Generation Through Trajectory-Distribution AlignmentYanxiao Sun, Jiafu Wu, Yun Cao, Chengming Xu et al.AAAI 2026 · 6 citations
- Towards One-step Causal Video Generation via Adversarial Self-DistillationYongqi Yang, Huayang Huang, Xu Peng, Xiaobin Hu et al.ICLR 2026 · 17 citations
- Scale-wise Distillation of Diffusion ModelsNikita Starodubcev, Ilya Drobyshevskiy, Denis Kuznedelev, Artem Babenko et al.ICLR 2026 · 13 citations
- OSV: One Step is Enough for High-Quality Image to Video GenerationXiaofeng Mao, Zhengkai Jiang, Fu-Yun Wang, Jiangning Zhang et al.CVPR 2025
