Masked Diffusion Models as Energy Minimization
Sitong Chen, Shen Nie, Jiacheng Sun, Zijin Feng, Zhenguo Li, Ji-Rong Wen, Chongxuan Li
Abstract
We present a systematic theoretical framework that interprets masked diffusion models (MDMs) as solutions to energy minimization problems in discrete optimal transport. Specifically, we prove that three distinct energy formulations--kinetic, conditional kinetic, and geodesic energy--are mathematically equivalent under the structure of MDMs, and that MDMs minimize all three when the mask schedule satisfies a closed-form optimality condition. This unification not only clarifies the theoretical foundations of MDMs, but also motivates practical improvements in sampling. By parameterizing interpolation schedules via Beta distributions, we reduce the schedule design space to a tractable 2D search, enabling efficient post-training tuning without model modification. Experiments on synthetic and real-world benchmarks demonstrate that our energy-inspired schedules outperform hand-crafted baselines, particularly in low-step sampling settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3767045-a6ab-4a46-9491-ec955d2cb4b4Cited by top-tier papers1
Ask how each one uses itBuilds on28
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
Related papers
- Score-Optimal Diffusion SchedulesChristopher Williams, Andrew Campbell, Arnaud Doucet, Saifuddin SyedNeurIPS 2024 · 19 citations
- Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference PoliciesChunsan Hong, Seonho An, Min-Soo Kim, Jong Chul YeICLR 2026 · 23 citations
- Align Your Steps: Optimizing Sampling Schedules in Diffusion ModelsAmirmojtaba Sabour, Sanja Fidler, Karsten KreisICML 2024 · 74 citations
- Conditional Diffusion SamplingFrancisco M Castro-Macías, Pablo Morales-Alvarez, Saifuddin Syed, Daniel Hernández-Lobato et al.ICML 2026 · 7 citations
- Flow Matching with General Discrete Paths: A Kinetic-Optimal PerspectiveNeta Shaul, Itai Gat, Marton Havasi, Daniel Severo et al.ICLR 2025
