InfoDLM: an Information-Adaptive Framework for Discrete Diffusion Language Model Pretraining
Shirou Jing, Chunshu Wu, Chuan Liu, Arghavan Bahadorinejad, Feitong Qiao, Dongfang Liu, Tony Geng
摘要
Diffusion language models (DLMs) can match or surpass similarly sized autoregressive language models on language understanding and reasoning. However, their mask-and-denoise pretraining relies on heuristic random masking, which fails to target the most informative tokens. Consequently, the model spends significant computational effort on redundant or trivial tokens. To address this, we propose InfoDLM, an adaptive DLM pretraining framework that reformulates mask selection as an active, feedback-driven process. InfoDLM targets tokens that offer the highest measurable information gain during mask selection. Specifically, we:
(1) introduce a Trainable Information-Gain (TIG) signal to quantify information gain of each masking configuration; (2) develop a feedback mechanism that adapts the masking policy to the model's evolving state with a maturity indicator; and (3) jointly optimize the DLM and masking policy through an interleaved training flow with only ∼14.1% wall-clock and ∼7.6% FLOPs overhead per cycle. Across reasoning-oriented benchmarks, InfoDLM achieves up to 13% improvement in reasoning accuracy over a small variant of LLaDA under comparable pretraining budgets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
- Improving Language Models by Retrieving from Trillions of TokensSebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai 等ICML 2022 · 被引用 1,629 次
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
相关 Paper
- Beyond Fully Random Masking: Attention-Guided Denoising and Optimization for Diffusion Language ModelsJia Deng, Junyi Li, Xin Zhao, Jinpeng Wang 等ACL 2026
- Learnability-Informed Fine-Tuning of Diffusion Language ModelsShubham Parashar, Atharv Chagi, Jacob Helwig, Lakshmi Madhavarapu 等ICML 2026
- Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked DiffusionsJaeyeon Kim, Kulin Shah, Vasilis Kontonis, Sham M. Kakade 等ICML 2025
- InforMask: Unsupervised Informative Masking for Language Model PretrainingNafis Sadeq, Canwen Xu, Julian J. McAuleyEMNLP 2022 · 被引用 12 次
- DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial AttentionYounjoo Lee, Seungkyun Dan, Junghoo Lee, Jaiyoung Park 等ICML 2026 · 被引用 2 次
