Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
Guanjie Chen, Shirui Huang, Yifu Sun, Kai Liu, Jianchen Zhu, Xiaoye Qu, Yu Cheng, Peng Chen
Abstract
Diffusion Models have emerged as a leading class of generative models, yet their iterative sampling process remains computationally expensive. Timestep distillation is a promising technique to accelerate generation, but it often requires extensive training and leads to image quality degradation. Furthermore, fine-tuning these distilled models for specific objectives, such as aesthetic appeal or user preference, using Reinforcement Learning (RL) is notoriously unstable and easily falls into reward hacking. In this work, we introduce Flash-DMD, a novel framework that enables fast convergence with distillation and joint RL-based refinement.Specifically, we first propose an efficient timestep-aware distillation strategy that significantly reduces training cost with enhanced realism, outperforming DMD2 with only its training cost. Second, we introduce a joint training scheme where the model is fine-tuned with an RL objective while the timestep distillation training continues simultaneously. We demonstrate that the stable, well-defined loss from the ongoing distillation acts as a powerful regularizer, effectively stabilizing the RL training process and preventing policy collapse. Extensive experiments on score-based and flow matching models show that our proposed Flash-DMD not only converges significantly faster but also achieves state-of-the-art generation quality in the few-step sampling regime, outperforming existing methods in visual quality, human preference, and text-image alignment metrics. Our work presents an effective paradigm for training efficient, high-fidelity, and stable generative models. Codes are attached in the supplementary.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- VA-π: Variational Policy Alignment for Pixel-Aware Autoregressive GenerationXinyao Liao, QIYUAN HE, Kai Xu, Xiaoye Qu et al.CVPR 2026 · 6 citations
- SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion DistillationZhuguanyu Wu, Ruihao Gong, Yang Yong, Yushi Huang et al.ICML 2026 · 2 citations
Builds on33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Distillation Models are Good Samplers for Diffusion Reinforcement LearningZunxu Liu, Aiqiu Wu, Zhaofan Qiu, Yingwei Pan et al.ICML 2026
- Simple and Fast Distillation of Diffusion ModelsZhenyu Zhou, Defang Chen, Can Wang, Chun Chen et al.NeurIPS 2024 · 44 citations
- One-Step Diffusion with Distribution Matching DistillationTianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman et al.CVPR 2024 · 75 citations
- Test-Time Reinforcement Learning for Flow MatchingJili Chen, Changqin Huang, Qionghao Huang, Yaxin Tu et al.ICML 2026
- DUO-VSR: Dual-Stream Distillation for One-Step Video Super-ResolutionZhengyao Lv, Menghan Xia, Xintao Wang, Kwan-Yee K. WongCVPR 2026 · 4 citations
