Theoretical Benefit and Limitation of Diffusion Language Model
Guhao Feng, Yihan Geng, Jian Guan, Wei Wu, Liwei Wang, Di He
摘要
Diffusion language models have emerged as a new approach for text generation. By enabling the parallel sampling of multiple tokens in each diffusion step, they appear to offer a more efficient alternative to auto-regressive models. However, our observations show that current open-sourced diffusion language models require more sampling steps to achieve comparable accuracy on representative tasks-resulting in even higher inference costs than their auto-regressive counterparts. To investigate whether this is an inherent limitation, we conduct a rigorous theoretical analysis of a widely adopted variant: the Masked Diffusion Model (MDM). Surprisingly, our analysis reveals that the conclusion is highly sensitive to the choice of evaluation metric. Under mild conditions, we prove that when the target is near-optimal perplexity, MDMs can achieve this goal in a constant number of sampling steps, independent of sequence length. This result demonstrates that efficiency can, in principle, be attained without compromising generation quality. However, when targeting low sequence error rate-which is important for assessing the "correctness" of a generated sequence, such as a reasoning chainwe show that in the worst case, the required sampling steps must scale linearly with sequence length, thereby eliminating the efficiency advantage. Our analysis establishes the first theoretical foundation for understanding the comparative strengths and limitations of MDMs, offering practical guidance on when to favor MDMs over auto-regressive models and vice versa.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- dKV-Cache: The Cache for Diffusion Language ModelsXinyin Ma, Runpeng Yu, Gongfan Fang, Xinchao WangNeurIPS 2025 · 被引用 145 次
- ParallelBench: Understanding the Trade-offs of Parallel Decoding in Diffusion LLMsWonjun Kang, Kevin Galim, Seunghyuk Oh, Minjae Lee 等ICLR 2026 · 被引用 49 次
- Multiverse: Your Language Models Secretly Decide How to Parallelize and Merge GenerationXinyu Yang, Yuwei An, Hongyi Liu, Tianqi Chen 等NeurIPS 2025 · 被引用 36 次
- MDNS: Masked Diffusion Neural Sampler via Stochastic Optimal ControlYuchen Zhu, Wei Guo, Jaemoo Choi, Guan-Horng Liu 等NeurIPS 2025 · 被引用 24 次
- Breaking AR's Sampling Bottleneck: Provable Acceleration via Diffusion Language ModelsGen Li, Changxiao CaiNeurIPS 2025 · 被引用 22 次
它引用的顶会 Paper34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
相关 Paper
- On Powerful Ways to Generate: Autoregression, Diffusion, and BeyondChenxiao Yang, Cai Zhou, David Wipf, Zhiyuan LiICLR 2026 · 被引用 7 次
- Diffusion Language Models are Provably Optimal Parallel SamplersHaozhe Jiang, Nika Haghtalab, Lijie ChenICLR 2026 · 被引用 5 次
- DiffuMamba: High-Throughput Diffusion LMs with Mamba BackboneVaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël, Eugene Belilovsky 等ICML 2026
- SPMDM: Enhancing Masked Diffusion Models through Simplifying Sampling PathYichen Zhu, Weiyu Chen, James Kwok, Zhou ZhaoNeurIPS 2025 · 被引用 1 次
- Scaling Beyond Masked Diffusion Language ModelsSubham Sekhar Sahoo, Jean-Marie Lemercier, Zhihan Yang, Justin Deschenaux 等ICML 2026 · 被引用 18 次
