Why Masking Diffusion Works: Condition on the Jump Schedule for Improved Discrete Diffusion
Alan Nawzad Amin, Nate Gruver, Andrew Gordon Wilson
Abstract
Discrete diffusion models, like continuous diffusion models, generate high-quality samples by gradually undoing noise applied to datapoints with a Markov process. Gradual generation in theory comes with many conceptual benefits; for example, inductive biases can be incorporated into the noising Markov process, and access to improved sampling algorithms. In practice, however, the consistently best performing discrete diffusion model is, surprisingly, masking diffusion, which does not denoise gradually. Here we explain the superior performance of masking diffusion by noting that it makes use of a fundamental difference between continuous and discrete Markov processes: discrete Markov processes evolve by discontinuous jumps at a fixed rate and, unlike other discrete diffusion models, masking diffusion builds in the known distribution of jump times and only learns where to jump to. We show that we can similarly bake in the known distribution of jump times into any discrete diffusion model. The resulting models - schedule-conditioned discrete diffusion (SCUD) - generalize classical discrete diffusion and masking diffusion. By applying SCUD to models with noising processes that incorporate inductive biases on images, text, and protein data, we build models that outperform masking.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Scaling Behavior of Discrete Diffusion Language ModelsDimitri von Rütte, Janis Fluri, Omead Pooladzandi, Bernhard Schölkopf et al.ICLR 2026 · 34 citations
- Coevolutionary Continuous Discrete Diffusion: Make Your Diffusion Language Model a Latent ReasonerCai Zhou, Chenxiao Yang, Yi Hu, Chenyu Wang et al.ICML 2026 · 21 citations
- Steering Generative Models with Experimental Data for Protein Fitness OptimizationJason Yang, Wenda Chu, Daniel Khalil, Raul Astudillo et al.NeurIPS 2025 · 13 citations
- Shrinking Proteins with DiffusionEthan Baron, Alan Nawzad Amin, Ruben Weitzman, Simon d'Oelsnitz et al.ICLR 2026 · 4 citations
- A Unification of Discrete, Gaussian, and Simplicial DiffusionNuria Alina Chandra, Yucen Lily Li, Alan Nawzad Amin, Alex Ali et al.ICLR 2026 · 2 citations
Builds on15
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Simplified and Generalized Masked Diffusion for Discrete DataJiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet et al.NeurIPS 2024 · 693 citations
- A Continuous Time Framework for Discrete Denoising ModelsAndrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth et al.NeurIPS 2022 · 496 citations
Related papers
- Is Your Diffusion Model Actually Denoising?Daniel Pfrommer, Zehao Dou, Christopher Scarvelis, Max Simchowitz et al.NeurIPS 2025 · 1 citation
- Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference PoliciesChunsan Hong, Seonho An, Min-Soo Kim, Jong Chul YeICLR 2026 · 23 citations
- Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion ProcessesBocheng Li, Zhujin Gao, Linli XuACL 2025
- Markup-to-Image Diffusion Models with Scheduled SamplingYuntian Deng, Noriyuki Kojima, Alexander M. RushICLR 2023 · 1 citation
- DiffusionBERT: Improving Generative Masked Language Models with Diffusion ModelsZhengfu He, Tianxiang Sun, Qiong Tang, Kuanning Wang et al.ACL 2023 · 63 citations
