PMI-Masking: Principled masking of correlated spans
Yoav Levine, Barak Lenz, Opher Lieber, Omri Abend, Kevin Leyton-Brown, Moshe Tennenholtz, Yoav Shoham
Abstract
Masking tokens uniformly at random constitutes a common flaw in the pretraining of Masked Language Models (MLMs) such as BERT. We show that such uniform masking allows an MLM to minimize its training objective by latching onto shallow local signals, leading to pretraining inefficiency and suboptimal downstream performance. To address this flaw, we propose PMI-Masking, a principled masking strategy based on the concept of Pointwise Mutual Information (PMI), which jointly masks a token n-gram if it exhibits high collocation over the corpus. PMI-Masking motivates, unifies, and improves upon prior more heuristic approaches that attempt to address the drawback of random uniform token masking, such as whole-word masking, entity/phrase masking, and random-span masking. Specifically, we show experimentally that PMI-Masking reaches the performance of prior masking approaches in half the training time, and consistently improves performance at the end of training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d98b09cb-0a5d-43b5-8ee3-a6cbca3262d1Cited by top-tier papers19
- PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document SummarizationWen Xiao, Iz Beltagy, Giuseppe Carenini, Arman CohanACL 2022 · 147 citations
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 119 citations
- Rethinking Tokenizer and Decoder in Masked Graph Modeling for MoleculesZhiyuan Liu, Yaorui Shi, An Zhang, Enzhi Zhang et al.NeurIPS 2023 · 71 citations
- Limits to Depth Efficiencies of Self-AttentionYoav Levine, Noam Wies, Or Sharir, Hofit Bata et al.NeurIPS 2020 · 60 citations
- 4M-21: An Any-to-Any Vision Model for Tens of Tasks and ModalitiesRoman Bachmann, Oguzhan Fatih Kar, David Mizrahi, Ali Garjani et al.NeurIPS 2024 · 60 citations
Builds on1
Related papers
- InforMask: Unsupervised Informative Masking for Language Model PretrainingNafis Sadeq, Canwen Xu, Julian J. McAuleyEMNLP 2022 · 12 citations
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang et al.ICML 2020 · 423 citations
- Learning Better Masking for Better Language Model Pre-trainingDongjie Yang, Zhuosheng Zhang, Hai ZhaoACL 2023 · 9 citations
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu et al.ACL 2022
- Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingMingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim et al.EMNLP 2022 · 8 citations
