Distilled Decoding 1: One-step Sampling of Image Auto-regressive Models with Flow Matching
Enshu Liu, Xuefei Ning, Yu Wang, Zinan Lin
Abstract
Image Auto-regressive (AR) models have emerged as a powerful paradigm of visual generative models. Despite their promising performance, they suffer from slow generation speed due to the large number of sampling steps required. Although Distilled Decoding 1 (DD1) was recently proposed to enable few-step sampling for image AR models, it still incurs significant performance degradation in the one-step setting, and relies on a pre-defined mapping that limits its flexibility. In this work, we propose a new method, Distilled Decoding 2 (DD2), to further advance the feasibility of one-step sampling for image AR models. Unlike DD1, DD2 does not without rely on a pre-defined mapping. We view the original AR model as a teacher model that provides the ground truth conditional score in the latent embedding space at each token position. Based on this, we propose a novel conditional score distillation loss to train a one-step generator. Specifically, we train a separate network to predict the conditional score of the generated distribution and apply score distillation at every token position conditioned on previous tokens. Experimental results show that DD2 enables one-step sampling for image AR models with a minimal FID increase from 3.40 to 5.43 and 4.11 to 7.58 on ImageNet-256, while achieving 8.0× and 238× speedup with VAR and LlamaGen models, respectively. Compared to the strongest baseline DD1, DD2 reduces the gap between the one-step sampling and original AR model by 67%, with up to 12.3× training speed-up simultaneously. DD2 takes a significant step toward the goal of one-step AR generation, opening up new possibilities for fast and high-quality AR modeling. Code is available at https://github.com/ imagination-research/Distilled-Decoding-2 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92a82ac1-30d2-4029-8eba-17c2f2ce8447Cited by top-tier papers11
- Synthesize Privacy-Preserving High-Resolution Images via Private Textual IntermediariesHaoxiang Wang, Zinan Lin, Da Yu, Huishuai ZhangNeurIPS 2025 · 9 citations
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image GenerationYao Teng, Fuyun Wang, Xian Liu, Zhekai Chen et al.NeurIPS 2025 · 8 citations
- Progressive Supernet Training for Efficient Visual Autoregressive ModelingXiaoyue Chen, Yuling Shi, Kaiyuan Li, Huandong Wang et al.CVPR 2026 · 7 citations
- Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft EmbeddingsYuanzhi Zhu, Xi Wang, Stéphane Lathuilière, Vicky KalogeitonICLR 2026 · 4 citations
- Distilled Decoding 2: One-step Sampling of Image Auto-regressive Models with Conditional Score DistillationEnshu Liu, Qian Chen, Xuefei Ning, Shengen Yan et al.NeurIPS 2025 · 4 citations
Builds on26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao et al.NeurIPS 2023 · 1,498 citations
Related papers
- Improved Distribution Matching Distillation for Fast Image SynthesisTianwei Yin, Michaël Gharbi, Taesung Park, Richard Zhang et al.NeurIPS 2024 · 728 citations
- One-Step Diffusion with Distribution Matching DistillationTianwei Yin, Michaël Gharbi, Richard Zhang, Eli Shechtman et al.CVPR 2024 · 75 citations
- One-Step Diffusion Distillation through Score Implicit MatchingWeijian Luo, Zemin Huang, Zhengyang Geng, J. Zico Kolter et al.NeurIPS 2024 · 81 citations
- DOLLAR: Few-Step Video Generation Via Distillation and Latent Reward OptimizationZihan Ding, Chi Jin, Difan Liu, Haitian Zheng et al.ICCV 2025 · 3 citations
- SwiftBrush: One-Step Text-to-Image Diffusion Model with Variational Score DistillationThuan Hoang Nguyen, Anh TranCVPR 2024 · 20 citations
