Masks Can Be Distracting: On Context Comprehension in Diffusion Language Models
Julianna Piskorz, Cristina Pinneri, Alvaro Correia, Motasem Alfarra, Risheek Garrepalli, Christos Louizos
摘要
Masked Diffusion Language Models (MDLMs) have recently emerged as a promising alternative to Autoregressive Language Models (ARLMs), leveraging a denoising objective that, in principle, should enable more uniform context utilisation. In this work, we examine the context comprehension abilities of MDLMs and uncover two key limitations. First, despite their more global training objective and bidirectional attention mechanism, similarly to ARLMS, MDLMs exhibit a strong locality bias : performance is highly sensitive to the position of relevant information within the input, favouring local over distant context. Second, appending a large number of mask tokens—required for generation—can significantly degrade context comprehension in models trained from scratch. Through systematic ablations, we find that these masks act as distractors , reducing the model's ability to process relevant information. To address and further study this undesirable behaviour, we introduce the mask-agnostic loss function that encourages predictions to remain invariant to the number of appended masks. Fine-tuning with this objective substantially mitigates the distracting effect of masks, improving robustness of MDLMs. Overall, our findings reveal critical limitations of the current MDLM training paradigm, with implications for training, evaluation and deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
- Discrete Diffusion Modeling by Estimating the Ratios of the Data DistributionAaron Lou, Chenlin Meng, Stefano ErmonICML 2024 · 被引用 473 次
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel DecodingChengyue Wu, Hao Zhang, Shuchen Xue, Zhijian Liu 等ICLR 2026 · 被引用 428 次
- Function Vectors in Large Language ModelsEric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller 等ICLR 2024 · 被引用 229 次
相关 Paper
- Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language ModelsSujung Hong, Chanyong Yoon, Seong Jae HwangICML 2026
- Unifying Masked Diffusion Models with Various Generation Orders and BeyondChunsan Hong, Sanghyun Lee, Jong Chul YEICML 2026
- Dual-objective Language Models: Training Efficiency Without OverfittingDavid Samuel, Lucas Georges Gabriel CharpentierICLR 2026
- Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and ArchitectureShuchen Xue, Tianyu Xie, Tianyang Hu, Zijin Feng 等ICML 2026 · 被引用 15 次
- Diffusion Beats Autoregressive in Data-Constrained SettingsMihir Prabhudesai, Mengning Wu, Amir Zadeh, Katerina Fragkiadaki 等NeurIPS 2025 · 被引用 69 次
