Plug-and-Play Diffusion Distillation
Yi-Ting Hsiao, Siavash Khodadadeh, Kevin Duarte, Wei-An Lin, Hui Qu, Mingi Kwon, Ratheesh Kalarot
Abstract
Diffusion models have shown tremendous results in image generation. However, due to the iterative nature of the diffusion process and its reliance on classifier-free guidance, inference times are slow. In this paper, we propose a new distillation approach for guided diffusion models in which an external lightweight guide model is trained while the original text-to-image model remains frozen.We show that our method reduces the inference computation of classifier-free guided latent-space diffusion models by almost half, and only requires 1% trainable parameters of the base model. Furthermore, once trained, our guide model can be applied to various fine-tuned, domain-specific versions of the base diffusion model without the need for additional training: this "plug-and-play" functionality drastically improves inference computation while maintaining the visual fidelity of generated images. Empirically, we show that our approach is able to produce visually appealing results and achieve a comparable FID score to the teacher with as few as 8 to 16 steps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9b42bd8-6181-4a11-9b3d-44b6d598e77dCited by top-tier papers12
- Flash Diffusion: Accelerating Any Conditional Diffusion Model for Few Steps Image GenerationClément Chadebec, Onur Tasar, Eyal Benaroche, Benjamin AubinAAAI 2025 · 52 citations
- Learnable Sparsity for Vision Generative ModelsYang Zhang, Er Jin, Wenzhong Liang, Yanfei Dong et al.ICLR 2026 · 8 citations
- Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language ModelsYimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof CzarneckiNeurIPS 2025 · 6 citations
- LBM: Latent Bridge Matching for Fast Image-to-Image TranslationClément Chadebec, Onur Tasar, Sanjeev Sreetharan, Benjamin AubinICCV 2025 · 5 citations
- DogFit: Domain-guided Fine-tuning for Efficient Transfer Learning of Diffusion ModelsYara Bahram, Mohammadhadi Shateri, Eric GrangerAAAI 2026 · 4 citations
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- Adaptive Guidance: Training-free Acceleration of Conditional Diffusion ModelsAngela Castillo, Jonas Kohler, Juan C. Pérez, Juan Pablo Pérez et al.AAAI 2025 · 1 citation
- SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two SecondsYanyu Li, Huan Wang, Qing Jin, Ju Hu et al.NeurIPS 2023 · 300 citations
- On Distillation of Guided Diffusion ModelsChenlin Meng, Robin Rombach, Ruiqi Gao, Diederik P. Kingma et al.CVPR 2023
- DICE: Distilling Classifier-Free Guidance into Text EmbeddingsZhenyu Zhou, Defang Chen, Can Wang, Chun Chen et al.AAAI 2026 · 2 citations
- Noise-free Score DistillationOren Katzir, Or Patashnik, Daniel Cohen-Or, Dani LischinskiICLR 2024 · 101 citations
