Breaking AR's Sampling Bottleneck: Provable Acceleration via Diffusion Language Models
Gen Li, Changxiao Cai
Abstract
Diffusion models have emerged as a powerful paradigm for modern generative modeling, demonstrating strong potential for large language models (LLMs). Unlike conventional autoregressive (AR) models that generate tokens sequentially, diffusion models allow for parallel sampling, offering a promising path to accelerate generation and eliminate the left-to-right generation constraints. Despite their empirical success, theoretical understandings of diffusion language models remain underdeveloped. In this work, we develop convergence guarantees for diffusion language models from an information-theoretic perspective. Our analysis demonstrates that the sampling error, measured by the Kullback-Leibler (KL) divergence, decays inversely with the number of iterations and scales linearly with the mutual information between tokens in the target text sequence. Crucially, our theory covers the regime , where is the text sequence length. This justifies that high-quality samples can be generated with fewer iterations than , thereby breaking the fundamental sampling bottleneck of steps required by AR models. We further establish matching upper and lower bounds, up to some constant factor, that shows the tightness of our convergence analysis. These results offer novel theoretical insights into the practical effectiveness of diffusion language models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03114638-e17e-4617-a8ba-40c70a3fb72eCited by top-tier papers3
- Learnable Sampler Distillation for Discrete Diffusion ModelsFeiyang Fu, Tongxian Guo, Zhaoqiang LiuNeurIPS 2025 · 10 citations
- Diffusion Models Are Statistically Optimal for Learning Low-Dimensional Multi-Modal DistributionsJingda Wu, Changxiao CaiICML 2026 · 1 citation
- Infinite Mask Diffusion for Few-Step DistillationJaehoon Yoo, Wonjung Kim, Chanhyuk Lee, Seunghoon HongICML 2026
Builds on33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
Related papers
- Theoretical Benefit and Limitation of Diffusion Language ModelGuhao Feng, Yihan Geng, Jian Guan, Wei Wu et al.NeurIPS 2025 · 52 citations
- Beyond Autoregression: Fast LLMs via Self-Distillation Through TimeJustin Deschenaux, Caglar GulcehreICLR 2025
- Accelerating Diffusion LLMs via Adaptive Parallel DecodingDaniel Israel, Guy Van den Broeck, Aditya GroverNeurIPS 2025 · 114 citations
- ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMsWonjun Kang, Kevin Galim, Seunghyuk Oh, Minjae Lee et al.ICLR 2026 · 49 citations
- dParallel: Learnable Parallel Decoding for dLLMsZigeng Chen, Gongfan Fang, Xinyin Ma, Ruonan Yu et al.ICLR 2026 · 64 citations
