InforMask: Unsupervised Informative Masking for Language Model Pretraining
Nafis Sadeq, Canwen Xu, Julian J. McAuley
Abstract
Masked language modeling is widely used for pretraining large language models for natural language understanding (NLU). However, random masking is suboptimal, allocating an equal masking rate for all tokens. In this paper, we propose InforMask, a new unsupervised masking strategy for training masked language models. InforMask exploits Pointwise Mutual Information (PMI) to select the most informative tokens to mask. We further propose two optimizations for InforMask to improve its efficiency. With a one-off preprocessing step, InforMask outperforms random masking and previously proposed masking strategies on the factual recall benchmark LAMA and the question answering benchmark SQuAD v1 and v2. 1 * Equal contribution. 1 The code and model checkpoints are available at https: //github.com/NafisSadeq/InforMask .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a635f096-c49a-4de2-a298-b943235e7fd4Cited by top-tier papers4
- DALE: Generative Data Augmentation for Low-Resource Legal NLPSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S. et al.EMNLP 2023 · 10 citations
- Understanding and Enhancing Mask-Based Pretraining towards Universal RepresentationsMingze Dong, Leda Wang, Yuval KlugerNeurIPS 2025 · 3 citations
- Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and ChallengeXuebo Liu, Yutong Wang, Derek F. Wong, Runzhe Zhan et al.ACL 2023 · 3 citations
- EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation LearningAshish Seth, Ramaneswaran Selvakumar, S. Sakshi, Sonal Kumar et al.EMNLP 2024 · 1 citation
Builds on10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- K-BERT: Enabling Language Representation with Knowledge GraphWeijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang et al.AAAI 2020 · 898 citations
Related papers
- PMI-Masking: Principled masking of correlated spansYoav Levine, Barak Lenz, Opher Lieber, Omri Abend et al.ICLR 2021 · 83 citations
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- InfoDLM: an Information-Adaptive Framework for Discrete Diffusion Language Model PretrainingShirou Jing, Chunshu Wu, Chuan Liu, Arghavan Bahadorinejad et al.ICML 2026
- Learning Better Masking for Better Language Model Pre-trainingDongjie Yang, Zhuosheng Zhang, Hai ZhaoACL 2023 · 9 citations
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderYi Liao, Xin Jiang, Qun LiuACL 2020 · 28 citations
