Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion Training
Jaeyeon Kim, Jonathan Geuter, David Alvarez-Melis, Sham Kakade, Sitan Chen
摘要
Masked Diffusion Models (MDMs) have emerged as a promising approach for generative modeling in discrete spaces. By generating sequences in any order and allowing for parallel decoding, they enable fast inference and strong performance on non-causal tasks. However, this flexibility comes with a training complexity trade-off: MDMs train on an exponentially large set of masking patterns, which is not only computationally expensive, but also creates a train--test mismatch between the random masks used in training and the highly structured masks induced by inference-time unmasking. In this work, we propose Progressive UnMAsking (PUMA), a simple modification of the forward masking process that aligns training-time and inference-time masking patterns, thereby focusing optimization on inference-aligned masks and speeding up training. Empirically, PUMA speeds up pretraining at the 125M scale by and offers complementary advantages on top of common recipes like autoregressive initialization. We open-source our codebase at https://github.com/JaeyeonKim01/PUMA.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible DecodingMarianne Arriola, Volodymyr KuleshovICML 2026 · 被引用 2 次
- Provable Sample Efficiency of Curriculum Post-Training for Transformer ReasoningDake Bu, Wei Huang, Andi Han, Atsushi Nitanda 等ICML 2026
它引用的顶会 Paper25
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
- Simplified and Generalized Masked Diffusion for Discrete DataJiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet 等NeurIPS 2024 · 被引用 693 次
相关 Paper
- Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked DiffusionsJaeyeon Kim, Kulin Shah, Vasilis Kontonis, Sham M. Kakade 等ICML 2025
- Unifying Masked Diffusion Models with Various Generation Orders and BeyondChunsan Hong, Sanghyun Lee, Jong Chul YEICML 2026
- Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical SamplingKaiwen Zheng, Yongxin Chen, Hanzi Mao, Ming-Yu Liu 等ICLR 2025
- Self-Speculative Masked DiffusionsAndrew Campbell, Valentin De Bortoli, Jiaxin Shi, Arnaud DoucetICLR 2026 · 被引用 12 次
- Any-Order Flexible Length Masked DiffusionJaeyeon Kim, Cheuk Lee Kit, Carles Domingo-Enrich, Yilun Du 等ICLR 2026 · 被引用 51 次
