Scaling Beyond Masked Diffusion Language Models
Subham Sekhar Sahoo, Jean-Marie Lemercier, Zhihan Yang, Justin Deschenaux, Jingyu Liu, John Thickstun, Ante Jukić
摘要
Diffusion language models are a promising alternative to autoregressive models due to their potential for faster generation. Among discrete diffusion approaches, Masked diffusion currently dominates, largely driven by strong perplexity on language modeling benchmarks. In this work, we present the first scaling law study of uniform-state and interpolating discrete diffusion methods. We also show that Masked diffusion models can be made approximately 12% more FLOPs-efficient when trained with a simple cross-entropy objective. We find that perplexity is informative within a diffusion family but can be misleading across families, where models with worse likelihood scaling may be preferable due to faster and more practical sampling, as reflected by the speedquality Pareto frontier. These results challenge the view that Masked diffusion is categorically the future of diffusion language modeling and that perplexity alone suffices for cross-algorithm comparison. Scaling all methods to 1.7B parameters, we show that uniform-state diffusion remains competitive on likelihood-based benchmarks and outperforms autoregressive and Masked diffusion models on GSM8K, despite worse validation perplexity. We provide the code, model checkpoints, and video tutorials on the project page: https://s-sahoo.com/scaling-dllms
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- The Diffusion Duality, Chapter II: ψ-Samplers and Efficient CurriculumJustin Deschenaux, Caglar Gulcehre, Subham Sekhar SahooICLR 2026 · 被引用 14 次
- Generalized Discrete Diffusion with Self-CorrectionLinxuan Wang, Ziyi Wang, Yikun Bai, Wei Deng 等ICML 2026 · 被引用 1 次
- Esoteric Language Models: A Family of Any-Order Diffusion LLMsSubham Sekhar Sahoo, Zhihan Yang, Yash Akhauri, Johnna Liu 等ICML 2026
它引用的顶会 Paper19
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
相关 Paper
- Scaling Behavior of Discrete Diffusion Language ModelsDimitri von Rütte, Janis Fluri, Omead Pooladzandi, Bernhard Schölkopf 等ICLR 2026 · 被引用 34 次
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
- Theoretical Benefit and Limitation of Diffusion Language ModelGuhao Feng, Yihan Geng, Jian Guan, Wei Wu 等NeurIPS 2025 · 被引用 52 次
- Scaling up Masked Diffusion Models on TextShen Nie, Fengqi Zhu, Chao Du, Tianyu Pang 等ICLR 2025
- Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and ArchitectureShuchen Xue, Tianyu Xie, Tianyang Hu, Zijin Feng 等ICML 2026 · 被引用 15 次
