Soft-Masked Diffusion Language Models
Michael Hersche, Samuel Moor-Smith, Thomas Hofmann, Abbas Rahimi
Abstract
Diffusion models have demonstrated strong potential in language modeling, offering various advantages over traditional autoregressive approaches. Their ability to generate and revise entire responses in parallel enables faster generation and built-in self-correction mechanisms. Most modern diffusion-based language models employ masked diffusion, where decoding involves iteratively processing masked tokens based on a binary decision: either retaining the mask or replacing it with the predicted token. However, this binary choice discards valuable predictive information when the mask is retained. To address this limitation, we introduce soft-masking (SM), a novel method that dynamically blends the embedding of the mask token with the embeddings of the top-k predicted tokens from the previous decoding step, for each retained mask. This provides the model with a more informative prior, preserving context from earlier computations and allowing partial information about masked tokens to propagate beyond a single step. We propose a training methodology that efficiently adapts masked diffusion language models to incorporate SM. We demonstrate that training a 169M parameter model from scratch with SM yields superior perplexity and MAUVE scores compared to binary masking baselines. Similarly, a pretrained model can be enhanced with SM through continued pretraining. Finally, we finetune two state-of-the-art diffusion models, Dream-7B and Dream-Coder-7B, with SM. SM consistently improves performance across multiple coding benchmarks, particularly in high-throughput settings. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 138ae4a6-46d4-4b87-9254-a57c033cd076Cited by top-tier papers3
- Beyond Hard Masks: Progressive Token Evolution for Diffusion Language ModelsLinhao Zhong, Linyu Wu, Bozhen Fang, Tianjian Feng et al.ACL 2026 · 4 citations
- Locally Coherent Parallel Decoding in Diffusion Language ModelsMichael Hersche, Nicolas Menet, Ronan Tanios, Abbas RahimiICML 2026 · 1 citation
- Edit-Based Refinement for Parallel Masked Diffusion Language ModelsHouxing Ren, Mingjie Zhan, Zimu Lu, Ke Wang et al.ICML 2026
Builds on44
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
Related papers
- Remasking Discrete Diffusion Models with Inference-Time ScalingGuanghan Wang, Yair Schiff, Subham S. Sahoo, Volodymyr KuleshovNeurIPS 2025 · 199 citations
- CORE: Context-Robust Remasking for Diffusion Language ModelsKevin Zhai, Sabbir Mollah, Zhenyi Wang, Mubarak ShahICML 2026 · 10 citations
- DreamOn: Diffusion Language Models For Code Infilling Beyond Fixed-size CanvasZirui Wu, Lin Zheng, Zhihui Xie, Jiacheng Ye et al.ICLR 2026 · 33 citations
- DiffusionBERT: Improving Generative Masked Language Models with Diffusion ModelsZhengfu He, Tianxiang Sun, Qiong Tang, Kuanning Wang et al.ACL 2023 · 63 citations
- Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language ModelsZemin Huang, Yuhang Wang, Zhiyang Chen, Guo-Jun QiICLR 2026 · 40 citations
