Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture
Shuchen Xue, Tianyu Xie, Tianyang Hu, Zijin Feng, Jiacheng Sun, Kenji Kawaguchi, Zhenguo Li, Zhi-Ming Ma
Abstract
Large language models (LLMs) predominantly use autoregressive (AR) approaches, but masked diffusion models (MDMs) are emerging as viable alternatives. A key challenge in comparing AR and MDM paradigms is their typical architectural difference: AR models are often decoder-only, while MDMs have largely been encoder-only. This practice of changing both the modeling paradigm and architecture simultaneously makes direct comparisons unfair, as it's hard to distinguish whether observed differences stem from the paradigm itself or the architectural shift. This research evaluates MDMs within a decoder-only framework to: (1) equitably compare MDM (as Any-Order AR, or AO-AR) and standard AR paradigms. Our investigation suggests that the standard AO-AR objective, which averages over all token permutations, may benefit from refinement, as many permutations appear less informative compared to the language's inherent left-to-right structure. (2) Investigate architectural influences (decoder-only vs. encoder-only) within MDMs. We demonstrate that while encoder-only MDMs model a simpler conditional probability space, decoder-only MDMs can achieve dramatic generation speedups (∼ 25×) and comparable perplexity with temperature annealing despite modeling a vastly larger space, highlighting key trade-offs. This work thus decouples core paradigm differences from architectural influences, offering insights for future model design. Code is available at https://github.com/scxue/AO-GPT-MDM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d591c6f5-c9a8-4620-8e7f-e3ca4d841e9dCited by top-tier papers6
- Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in SpeedYonggan Fu, Lexington Whalen, Zhifan Ye, Xin Dong et al.ICML 2026 · 22 citations
- On Powerful Ways to Generate: Autoregression, Diffusion, and BeyondChenxiao Yang, Cai Zhou, David Wipf, Zhiyuan LiICLR 2026 · 7 citations
- Masks Can Be Distracting: On Context Comprehension in Diffusion Language ModelsJulianna Piskorz, Cristina Pinneri, Alvaro Correia, Motasem Alfarra et al.ICML 2026 · 5 citations
- DiffGRM: Diffusion-based Generative Recommendation ModelZhao Liu, Yichen Zhu, Yiqing Yang, Xiao Lv et al.WWW 2026 · 2 citations
- Dual-objective Language Models: Training Efficiency Without OverfittingDavid Samuel, Lucas Georges Gabriel CharpentierICLR 2026
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
Related papers
- Unifying Masked Diffusion Models with Various Generation Orders and BeyondChunsan Hong, Sanghyun Lee, Jong Chul YEICML 2026
- Diffusion Beats Autoregressive in Data-Constrained SettingsMihir Prabhudesai, Mengning Wu, Amir Zadeh, Katerina Fragkiadaki et al.NeurIPS 2025 · 69 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Scaling Beyond Masked Diffusion Language ModelsSubham Sekhar Sahoo, Jean-Marie Lemercier, Zhihan Yang, Justin Deschenaux et al.ICML 2026 · 18 citations
- The Efficiency Gap in Byte ModelingCeline Lee, Jing Nathan Yan, Chen Liang, Jiaxin Shi et al.ICML 2026
