Learning Better Masking for Better Language Model Pre-training
Dongjie Yang, Zhuosheng Zhang, Hai Zhao
摘要
Masked Language Modeling (MLM) has been widely used as the denoising objective in pretraining language models (PrLMs). Existing PrLMs commonly adopt a Random-Token Masking strategy where a fixed masking ratio is applied and different contents are masked by an equal probability throughout the entire training. However, the model may receive a complicated impact from pre-training status, which changes accordingly as training time goes on. In this paper, we show that such time-invariant MLM settings on masking ratio and masked content are unlikely to deliver an optimal outcome, which motivates us to explore the influence of time-variant MLM settings. We propose two scheduled masking approaches that adaptively tune the masking ratio and masked content in different training stages, which improves the pre-training efficiency and effectiveness verified on the downstream tasks. Our work is a pioneer study on time-variant masking strategy on ratio and content and gives a better understanding of how masking ratio and masked content influence the MLM pretraining 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- KidLM: Advancing Language Models for Children - Early Insights and Future DirectionsMir Tafseer Nayeem, Davood RafieiEMNLP 2024 · 被引用 7 次
- Reformulating NLP tasks to Capture Longitudinal Manifestation of Language Disorders in People with DementiaDimitris Gkoumas, Matthew Purver, Maria LiakataEMNLP 2023 · 被引用 4 次
- Improving Pretraining Techniques for Code-Switched NLPRicheek Das, Sahasra Ranjan, Shreya Pathak, Preethi JyothiACL 2023 · 被引用 3 次
- Understanding and Enhancing Mask-Based Pretraining towards Universal RepresentationsMingze Dong, Leda Wang, Yuval KlugerNeurIPS 2025 · 被引用 3 次
- A Content-Preserving Secure Linguistic SteganographyLingyun Xiang, Chengfu Ou, Xu He, Zhongliang Yang 等AAAI 2026
它引用的顶会 Paper5
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-TrainingHangbo Bao, Li Dong, Furu Wei, Wenhui Wang 等ICML 2020 · 被引用 423 次
- PMI-Masking: Principled masking of correlated spansYoav Levine, Barak Lenz, Opher Lieber, Omri Abend 等ICLR 2021 · 被引用 83 次
- Pre-training Universal Language RepresentationYian Li, Hai ZhaoACL 2021
相关 Paper
- InforMask: Unsupervised Informative Masking for Language Model PretrainingNafis Sadeq, Canwen Xu, Julian J. McAuleyEMNLP 2022 · 被引用 12 次
- On the Influence of Masking Policies in Intermediate Pre-trainingQinyuan Ye, Belinda Z. Li, Sinong Wang, Benjamin Bolte 等EMNLP 2021 · 被引用 2 次
- Efficient Pre-training of Masked Language Model via Concept-based Curriculum MaskingMingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim 等EMNLP 2022 · 被引用 8 次
- Probabilistically Masked Language Model Capable of Autoregressive Generation in Arbitrary Word OrderYi Liao, Xin Jiang, Qun LiuACL 2020 · 被引用 28 次
- DiffusionBERT: Improving Generative Masked Language Models with Diffusion ModelsZhengfu He, Tianxiang Sun, Qiong Tang, Kuanning Wang 等ACL 2023 · 被引用 63 次
