Learning Unmasking Policies for Diffusion Language Models
Metod Jazbec, Theo X. Olausson, Louis Béthune, Pierre Ablin, Michael Kirchhof, Joao Monteiro, Victor Guilherme Turrisi da Costa, Jason Ramapuram, Marco Cuturi
Abstract
Diffusion (Large) Language Models (dLLMs) now match the downstream performance of their autoregressive counterparts on many tasks, while holding the promise of being more efficient during inference. One critical design aspect of dLLMs is the sampling procedure that selects which tokens to unmask at each diffusion step. Indeed, recent work has found that heuristic strategies such as confidence thresholding improve both sample quality and token throughput compared to random unmasking. However, such heuristics have downsides: they require manual tuning, and we observe that their performance degrades with larger block sizes. In this work, we instead propose to train sampling procedures using reinforcement learning. Specifically, we formalize masked diffusion sampling as a Markov decision process in which the dLLM serves as the environment, and propose a lightweight policy based on a single-layer transformer that maps dLLM token confidences to unmasking decisions. Our experiments show that these trained policies match the performance of state-of-the-art heuristics when combined with semi-autoregressive (block) generation, while outperforming them in the full-diffusion setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion TrainingJaeyeon Kim, Jonathan Geuter, David Alvarez-Melis, Sham Kakade et al.ICML 2026 · 8 citations
- Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from -ParityJianhao Huang, Baharan MirzasoleimanICML 2026 · 2 citations
- DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial AttentionYounjoo Lee, Seungkyun Dan, Junghoo Lee, Jaiyoung Park et al.ICML 2026 · 2 citations
- Scheduling Thoughts: Learning the Order of Thought in Diffusion Language ModelsJiawei Xu, Minghui Liu, Aakriti Agrawal, Yifan Chen et al.ICML 2026 · 1 citation
- LUGS: Latent-aware Guidance for Efficient Unmasking in Diffusion Large Language ModelsNuanqiao Shan, Kairong Han, Xinpeng Dong, Kun KuangICML 2026
Builds on36
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Argmax Flows and Multinomial Diffusion: Learning Categorical DistributionsEmiel Hoogeboom, Didrik Nielsen, Priyank Jaini, Patrick Forré et al.NeurIPS 2021 · 782 citations
- Simplified and Generalized Masked Diffusion for Discrete DataJiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet et al.NeurIPS 2024 · 693 citations
Related papers
- TA-GRPO-d: Trajectory-Aware GRPO for Optimizing Denoising Trajectories in Diffusion LLMsGyunyeop Kim, Sangwoo KangACL 2026
- MRO: Enhancing Reasoning in Diffusion Language Models via Multi-Reward OptimizationChenglong Wang, Yang Gan, Hang Zhou, Chi Hu et al.NeurIPS 2025 · 4 citations
- Reward-Weighted Sampling: Enhancing Non-Autoregressive Characteristics in Masked Diffusion LLMsDaehoon Gwak, Minseo Jung, Junwoo Park, Minho Park et al.EMNLP 2025
- Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language ModelsZemin Huang, Yuhang Wang, Zhiyang Chen, Guo-Jun QiICLR 2026 · 40 citations
- Diffusion Language Models are Provably Optimal Parallel SamplersHaozhe Jiang, Nika Haghtalab, Lijie ChenICLR 2026 · 5 citations
