One-Step Diffusion Distillation via Deep Equilibrium Models
Zhengyang Geng, Ashwini Pokle, J. Zico Kolter
Abstract
Diffusion models excel at producing high-quality samples but naively require hundreds of iterations, prompting multiple attempts to distill the generation process into a faster network. However, many existing approaches suffer from a variety of challenges: the process for distillation training can be complex, often requiring multiple training stages, and the resulting models perform poorly when utilized in single-step generative applications. In this paper, we introduce a simple yet effective means of distilling diffusion models directly from initial noise to the resulting image. Of particular importance to our approach is to leverage a new Deep Equilibrium (DEQ) model as the distilled architecture: the Generative Equilibrium Transformer (GET). Our method enables fully offline training with just noise/image pairs from the diffusion model while achieving superior performance compared to existing one-step methods on comparable training budgets. We demonstrate that the DEQ architecture is crucial to this capability, as GET matches a larger ViT in terms of FID scores while striking a critical balance of computational cost and image quality. Code, checkpoints, and datasets are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 31804e7b-9e93-4a8d-bd0c-bdc2d8d86f3fCited by top-tier papers44
- Mean Flows for One-step Generative ModelingZhengyang Geng, Mingyang Deng, Xingjian Bai, Zico Kolter et al.NeurIPS 2025 · 628 citations
- Score identity Distillation: Exponentially Fast Distillation of Pretrained Diffusion Models for One-Step GenerationMingyuan Zhou, Huangjie Zheng, Zhendong Wang, Mingzhang Yin et al.ICML 2024 · 174 citations
- Phased Consistency ModelsFu-Yun Wang, Zhaoyang Huang, Alexander William Bergman, Dazhong Shen et al.NeurIPS 2024 · 86 citations
- One-Step Diffusion Distillation through Score Implicit MatchingWeijian Luo, Zemin Huang, Zhengyang Geng, J. Zico Kolter et al.NeurIPS 2024 · 81 citations
- Utilizing Image Transforms and Diffusion Models for Generative Modeling of Short and Long Time SeriesIlan Naiman, Nimrod Berman, Itai Pemper, Idan Arbiv et al.NeurIPS 2024 · 69 citations
Builds on61
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- Data-free Distillation of Diffusion Models with BootstrappingJiatao Gu, Chen Wang, Shuangfei Zhai, Yizhe Zhang et al.ICML 2024 · 5 citations
- One-Step Diffusion with Distribution Matching DistillationTianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman et al.CVPR 2024 · 75 citations
- Progressive Distillation for Fast Sampling of Diffusion ModelsTim Salimans, Jonathan HoICLR 2022 · 9 citations
- Simple Distillation for One-Step Diffusion ModelsHuaisheng Zhu, Teng Xiao, Shijie Zhou, Zhimeng Guo et al.NeurIPS 2025 · 7 citations
- EM Distillation for One-step Diffusion ModelsSirui Xie, Zhisheng Xiao, Diederik P. Kingma, Tingbo Hou et al.NeurIPS 2024 · 69 citations
