Any-Order Flexible Length Masked Diffusion
Jaeyeon Kim, Cheuk Lee Kit, Carles Domingo-Enrich, Yilun Du, Sham M. Kakade, Timothy Ngotiaoco, Sitan Chen, Michael S. Albergo
Abstract
Masked diffusion models (MDMs) have recently emerged as a promising alternative to autoregressive models over discrete domains. MDMs generate sequences in an any-order, parallel fashion, enabling fast inference and strong performance on non-causal tasks. However, a crucial limitation is that they do not support token insertions and are thus limited to fixed-length generations. To this end, we introduce Flexible Masked Diffusion Models (FlexMDMs), a discrete diffusion paradigm that simultaneously can model sequences of flexible length while provably retaining MDMs' flexibility of any-order inference. Grounded in an extension of the stochastic interpolant framework, FlexMDMs generate sequences by inserting mask tokens and unmasking them. Empirically, we show that FlexMDMs match MDMs in perplexity while modeling length statistics with much higher fidelity. On a synthetic maze planning task, they achieve ≈ 60% higher success rate than MDM baselines. Finally, we show pretrained MDMs can easily be retrofitted into FlexMDMs: on 16 H100s, it takes only three days to fine-tune LLaDA-8B into a FlexMDM, achieving superior performance on math (GSM8K, 58%→67%) and code infilling performance (52%→65%). ⋆ equal contribution, randomized ordering; Lee Cheuk-Kit led the theoretical development of the method and engineering effort; Jaeyeon Kim led experiment development and paper presentation. † lead senior authors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- DreamOn: Diffusion Language Models For Code Infilling Beyond Fixed-size CanvasZirui Wu, Lin Zheng, Zhihui Xie, Jiacheng Ye et al.ICLR 2026 · 33 citations
- Rainbow Padding: Mitigating Early Termination in Instruction-Tuned Diffusion LLMsBumjun Kim, Dongjae Jeon, Dueun Kim, Wonje Jeung et al.ICLR 2026 · 11 citations
- CORE: Context-Robust Remasking for Diffusion Language ModelsKevin Zhai, Sabbir Mollah, Zhenyi Wang, Mubarak ShahICML 2026 · 10 citations
- Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion TrainingJaeyeon Kim, Jonathan Geuter, David Alvarez-Melis, Sham Kakade et al.ICML 2026 · 8 citations
- Beyond Masks: Efficient, Flexible Diffusion Language Models via Deletion-Insertion ProcessesFangyu Ding, Ding Ding, Sijin Chen, Kaibo Wang et al.ICLR 2026 · 7 citations
Builds on39
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
Related papers
- Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked DiffusionsJaeyeon Kim, Kulin Shah, Vasilis Kontonis, Sham M. Kakade et al.ICML 2025
- On Powerful Ways to Generate: Autoregression, Diffusion, and BeyondChenxiao Yang, Cai Zhou, David Wipf, Zhiyuan LiICLR 2026 · 7 citations
- SPMDM: Enhancing Masked Diffusion Models through Simplifying Sampling PathYichen Zhu, Weiyu Chen, James Kwok, Zhou ZhaoNeurIPS 2025 · 1 citation
- Insertion Based Sequence Generation with Learnable Order DynamicsDhruvesh Patel, Benjamin Rozonoyer, Gaurav Pandey, Tahira Naseem et al.ICML 2026 · 1 citation
- Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion ForcingXu Wang, Chenkai Xu, Yijie Jin, Jiachun Jin et al.ICLR 2026 · 116 citations
