On the Reasoning Abilities of Masked Diffusion Language Models
Anej Svete, Ashish Sabharwal
Abstract
Masked diffusion models (MDMs) for text offer a compelling alternative to traditional autoregressive language models. Parallel generation makes them efficient, but their computational capabilities and the limitations inherent in their parallelism remain largely unexplored. To this end, we characterize what types of reasoning problems MDMs can provably solve and how efficiently. We do this by connecting MDMs to the well-understood reasoning frameworks of chain of thought (CoT) and padded looped transformers (PLTs) in the finite-precision log-width setting: We show that MDMs and polynomially-padded PLTs are, in fact, equivalent in this setting, and that MDMs can solve all problems that CoT-augmented transformers can. Moreover, we showcase classes of problems (including regular languages) for which MDMs are inherently more efficient than CoT transformers, where parallel generation allows for substantially faster reasoning. ˚This research was conducted while interning at the Allen Institute for AI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4b173c85-d92c-41c9-95ca-934ffe038138Cited by top-tier papers2
- A Formal Comparison Between Chain of Thought and Latent ThoughtKevin Xu, Issei SatoICML 2026 · 12 citations
- Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don'tAnej Svete, William Merrill, Ryan Cotterell, Ashish SabharwalICML 2026 · 3 citations
Builds on21
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Discrete Diffusion Modeling by Estimating the Ratios of the Data DistributionAaron Lou, Chenlin Meng, Stefano ErmonICML 2024 · 473 citations
- Towards Revealing the Mystery behind Chain of Thought: A Theoretical PerspectiveGuhao Feng, Bohang Zhang, Yuntian Gu, Haotian Ye et al.NeurIPS 2023 · 470 citations
- Chain of Thought Empowers Transformers to Solve Inherently Serial ProblemsZhiyuan Liu, Hong Liu, Denny Zhou, Tengyu MaICLR 2024 · 259 citations
Related papers
- Diffusion Language Models are Provably Optimal Parallel SamplersHaozhe Jiang, Nika Haghtalab, Lijie ChenICLR 2026 · 5 citations
- Exact Expressive Power of Transformers with PaddingWilliam Merrill, Ashish SabharwalNeurIPS 2025 · 21 citations
- On Powerful Ways to Generate: Autoregression, Diffusion, and BeyondChenxiao Yang, Cai Zhou, David Wipf, Zhiyuan LiICLR 2026 · 7 citations
- Diffusion of Thought: Chain-of-Thought Reasoning in Diffusion Language ModelsJiacheng Ye, Shansan Gong, Liheng Chen, Lin Zheng et al.NeurIPS 2024 · 47 citations
- Do Efficient Transformers Really Save Computation?Kai Yang, Jan Ackermann, Zhenyu He, Guhao Feng et al.ICML 2024 · 29 citations
