Efficient Pre-training of Masked Language Model via Concept-based Curriculum Masking
Mingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim, SangKeun Lee
摘要
Self-supervised pre-training has achieved remarkable success in extensive natural language processing tasks. Masked language modeling (MLM) has been widely used for pre-training effective bidirectional representations but comes at a substantial training cost. In this paper, we propose a novel concept-based curriculum masking (CCM) method to efficiently pre-train a language model. CCM has two key differences from existing curriculum learning approaches to effectively reflect the nature of MLM. First, we introduce a novel curriculum that evaluates the MLM difficulty of each token based on a carefully-designed linguistic difficulty criterion. Second, we construct a curriculum that masks easy words and phrases first and gradually masks related ones to the previously masked ones based on a knowledge graph. Experimental results show that CCM significantly improves pre-training efficiency. Specifically, the model trained with CCM shows comparative performance with the original BERT on the General Language Understanding Evaluation benchmark at half of the training cost.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Improving Instruction Following in Language Models through Proxy-Based Uncertainty EstimationJoonHo Lee, Jae Oh Woo, Juree Seok, Parisa Hassanzadeh 等ICML 2024 · 被引用 4 次
- Empower Nested Boolean Logic via Self-Supervised Curriculum LearningHongqiu Wu, Linfeng Liu, Hai Zhao, Min ZhangEMNLP 2023 · 被引用 2 次
- InfoDLM: an Information-Adaptive Framework for Discrete Diffusion Language Model PretrainingShirou Jing, Chunshu Wu, Chuan Liu, Arghavan Bahadorinejad 等ICML 2026
- Curriculum Debiasing: Toward Robust Parameter-Efficient Fine-Tuning Against Dataset BiasesMingyu Lee, Yeachan Kim, Wing-Lam Mok, SangKeun LeeACL 2025
它引用的顶会 Paper8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 被引用 541 次
- ConvBERT: Improving BERT with Span-based Dynamic ConvolutionZihang Jiang, Weihao Yu, Daquan Zhou, Yunpeng Chen 等NeurIPS 2020 · 被引用 220 次
- Curriculum Learning for Natural Language UnderstandingBenfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang 等ACL 2020 · 被引用 156 次
相关 Paper
- Norm-Based Curriculum Learning for Neural Machine TranslationXuebo Liu, Houtim Lai, Derek F. Wong, Lidia S. ChaoACL 2020 · 被引用 97 次
- Self-supervised Masked Graph Autoencoder via Structure-aware CurriculumHaoyang Li, Xin Wang, Zeyang Zhang, Zongyuan Wu 等ICML 2025
- Exploiting Structured Knowledge in Text via Graph-Guided Representation LearningTao Shen, Yi Mao, Pengcheng He, Guodong Long 等EMNLP 2020 · 被引用 60 次
- PMI-Masking: Principled masking of correlated spansYoav Levine, Barak Lenz, Opher Lieber, Omri Abend 等ICLR 2021 · 被引用 83 次
- Pre-Training Curriculum for Multi-Token Prediction in Language ModelsAnsar Aynetdinov, Alan AkbikACL 2025
