PMI-Masking: Principled masking of correlated spans
Yoav Levine, Barak Lenz, Opher Lieber, Omri Abend, Kevin Leyton-Brown, Moshe Tennenholtz, Yoav Shoham
摘要
Masking tokens uniformly at random constitutes a common flaw in the pretraining of Masked Language Models (MLMs) such as BERT. We show that such uniform masking allows an MLM to minimize its training objective by latching onto shallow local signals, leading to pretraining inefficiency and suboptimal downstream performance. To address this flaw, we propose PMI-Masking, a principled masking strategy based on the concept of Pointwise Mutual Information (PMI), which jointly masks a token n-gram if it exhibits high collocation over the corpus. PMI-Masking motivates, unifies, and improves upon prior more heuristic approaches that attempt to address the drawback of random uniform token masking, such as whole-word masking, entity/phrase masking, and random-span masking. Specifically, we show experimentally that PMI-Masking reaches the performance of prior masking approaches in half the training time, and consistently improves performance at the end of training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- PRIMERA: Pyramid-based Masked Sentence Pre-training for Multi-document SummarizationWen Xiao, Iz Beltagy, Giuseppe Carenini, Arman CohanACL 2022 · 被引用 147 次
- Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt TuningColin Wei, Sang Michael Xie, Tengyu MaNeurIPS 2021 · 被引用 119 次
- Rethinking Tokenizer and Decoder in Masked Graph Modeling for MoleculesZhiyuan Liu, Yaorui Shi, An Zhang, Enzhi Zhang 等NeurIPS 2023 · 被引用 71 次
- Limits to Depth Efficiencies of Self-AttentionYoav Levine, Noam Wies, Or Sharir, Hofit Bata 等NeurIPS 2020 · 被引用 60 次
- 4M-21: An Any-to-Any Vision Model for Tens of Tasks and ModalitiesRoman Bachmann, Oguzhan Fatih Kar, David Mizrahi, Ali Garjani 等NeurIPS 2024 · 被引用 60 次
它引用的顶会 Paper1
相关 Paper
- InforMask: Unsupervised Informative Masking for Language Model PretrainingNafis Sadeq, Canwen Xu, Julian J. McAuleyEMNLP 2022 · 被引用 12 次
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
- Learning Better Masking for Better Language Model Pre-trainingDongjie Yang, Zhuosheng Zhang, Hai ZhaoACL 2023 · 被引用 9 次
- Token Dropping for Efficient BERT PretrainingLe Hou, Richard Yuanzhe Pang, Tianyi Zhou, Yuexin Wu 等ACL 2022
- Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingMingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim 等EMNLP 2022 · 被引用 8 次
