Lookahead Unmasking Elicits Reliable Decoding in Diffusion Language Models
Sanghyun Lee, Seungryong Kim, Jongho Park, Dongmin Park
Abstract
Masked Diffusion Models (MDMs) as language models generate by iteratively unmasking tokens, yet their performance crucially depends on the inference-time order of unmasking. Conventional methods such as confidence-based sampling are short-sighted, focusing on local optimization while neglecting additional testtime computation and allowing early decoding errors to cascade. We propose Lookahead Unmasking (LookUM), which addresses these concerns by reformulating sampling as path selection over all possible unmasking orders without the need for an external reward model. Our framework couples (i) a path generator that proposes paths by sampling from pools of unmasking sets with (ii) a verifier that computes the uncertainty of the proposed paths and performs importance sampling to subsequently select the final paths. Empirically, erroneous unmasking measurably inflates sequence-level uncertainty, and our method exploits this to avoid error-prone trajectories. We validate our framework across six benchmarks, such as mathematics, planning, and coding, and demonstrate consistent performance improvements. LookUM requires only two to three paths to achieve peak performance, demonstrating remarkably efficient path selection. The consistent improvements on both LLaDA and post-trained LLaDA 1.5 are particularly striking: base LLaDA with LookUM rivals the performance of RLtuned LLaDA 1.5, while LookUM further enhances LLaDA 1.5 itself-showing that uncertainty-based verification provides orthogonal benefits to reinforcement learning and underscoring the versatility of our framework.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
Related papers
- Lookahead Path Likelihood Optimization for Diffusion LLMsXuejie Liu, Vit Chun Yap, Yitao Liang, Anji LiuICML 2026 · 1 citation
- Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language ModelsZemin Huang, Zhiyang Chen, Zijun Wang, Tiancheng Li et al.NeurIPS 2025 · 55 citations
- Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked DiffusionsJaeyeon Kim, Kulin Shah, Vasilis Kontonis, Sham M. Kakade et al.ICML 2025
- Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference PoliciesChunsan Hong, Seonho An, Min-Soo Kim, Jong Chul YeICLR 2026 · 23 citations
- Improving Sampling for Masked Diffusion Models via Information GainKaisen Yang, Jayden Teoh, Kaicheng Yang, Yitong Zhang et al.ICML 2026 · 6 citations
