InforMask: Unsupervised Informative Masking for Language Model Pretraining
Nafis Sadeq, Canwen Xu, Julian J. McAuley
摘要
Masked language modeling is widely used for pretraining large language models for natural language understanding (NLU). However, random masking is suboptimal, allocating an equal masking rate for all tokens. In this paper, we propose InforMask, a new unsupervised masking strategy for training masked language models. InforMask exploits Pointwise Mutual Information (PMI) to select the most informative tokens to mask. We further propose two optimizations for InforMask to improve its efficiency. With a one-off preprocessing step, InforMask outperforms random masking and previously proposed masking strategies on the factual recall benchmark LAMA and the question answering benchmark SQuAD v1 and v2. 1 * Equal contribution. 1 The code and model checkpoints are available at https: //github.com/NafisSadeq/InforMask .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DALE: Generative Data Augmentation for Low-Resource Legal NLPSreyan Ghosh, Chandra Kiran Reddy Evuru, Sonal Kumar, Ramaneswaran S. 等EMNLP 2023 · 被引用 10 次
- Understanding and Enhancing Mask-Based Pretraining towards Universal RepresentationsMingze Dong, Leda Wang, Yuval KlugerNeurIPS 2025 · 被引用 3 次
- Revisiting Commonsense Reasoning in Machine Translation: Training, Evaluation and ChallengeXuebo Liu, Yutong Wang, Derek F. Wong, Runzhe Zhan 等ACL 2023 · 被引用 3 次
- EH-MAM: Easy-to-Hard Masked Acoustic Modeling for Self-Supervised Speech Representation LearningAshish Seth, Ramaneswaran Selvakumar, S. Sakshi, Sonal Kumar 等EMNLP 2024 · 被引用 1 次
它引用的顶会 Paper10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat 等ICML 2020 · 被引用 2,937 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- K-BERT: Enabling Language Representation with Knowledge GraphWeijie Liu, Peng Zhou, Zhe Zhao, Zhiruo Wang 等AAAI 2020 · 被引用 898 次
相关 Paper
- PMI-Masking: Principled masking of correlated spansYoav Levine, Barak Lenz, Opher Lieber, Omri Abend 等ICLR 2021 · 被引用 83 次
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
- InfoDLM: an Information-Adaptive Framework for Discrete Diffusion Language Model PretrainingShirou Jing, Chunshu Wu, Chuan Liu, Arghavan Bahadorinejad 等ICML 2026
- Learning Better Masking for Better Language Model Pre-trainingDongjie Yang, Zhuosheng Zhang, Hai ZhaoACL 2023 · 被引用 9 次
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderYi Liao, Xin Jiang, Qun LiuACL 2020 · 被引用 28 次
