Efficient Pre-training of Masked Language Model via Concept-based Curriculum Masking
Mingyu Lee, Jun-Hyung Park, Junho Kim, Kang-Min Kim, SangKeun Lee
Abstract
Self-supervised pre-training has achieved remarkable success in extensive natural language processing tasks. Masked language modeling (MLM) has been widely used for pre-training effective bidirectional representations but comes at a substantial training cost. In this paper, we propose a novel concept-based curriculum masking (CCM) method to efficiently pre-train a language model. CCM has two key differences from existing curriculum learning approaches to effectively reflect the nature of MLM. First, we introduce a novel curriculum that evaluates the MLM difficulty of each token based on a carefully-designed linguistic difficulty criterion. Second, we construct a curriculum that masks easy words and phrases first and gradually masks related ones to the previously masked ones based on a knowledge graph. Experimental results show that CCM significantly improves pre-training efficiency. Specifically, the model trained with CCM shows comparative performance with the original BERT on the General Language Understanding Evaluation benchmark at half of the training cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce8f6968-7ab8-468b-a805-d0b9d06e078dCited by top-tier papers4
- Improving Instruction Following in Language Models through Proxy-Based Uncertainty EstimationJoonHo Lee, Jae Oh Woo, Juree Seok, Parisa Hassanzadeh et al.ICML 2024 · 4 citations
- Empower Nested Boolean Logic via Self-Supervised Curriculum LearningHongqiu Wu, Linfeng Liu, Hai Zhao, Min ZhangEMNLP 2023 · 2 citations
- InfoDLM: an Information-Adaptive Framework for Discrete Diffusion Language Model PretrainingShirou Jing, Chunshu Wu, Chuan Liu, Arghavan Bahadorinejad et al.ICML 2026
- Curriculum Debiasing: Toward Robust Parameter-Efficient Fine-Tuning Against Dataset BiasesMingyu Lee, Yeachan Kim, Wing-Lam Mok, SangKeun LeeACL 2025
Builds on8
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- ELECTRA: Pre-training Text Encoders as Discriminators Rather Than GeneratorsKevin Clark, Minh-Thang Luong, Quoc V. Le, Christopher D. ManningICLR 2020 · 541 citations
- ConvBERT: Improving BERT with Span-based Dynamic ConvolutionZihang Jiang, Weihao Yu, Daquan Zhou, Yunpeng Chen et al.NeurIPS 2020 · 220 citations
- Curriculum Learning for Natural Language UnderstandingBenfeng Xu, Licheng Zhang, Zhendong Mao, Quan Wang et al.ACL 2020 · 156 citations
Related papers
- Norm-Based Curriculum Learning for Neural Machine TranslationXuebo Liu, Houtim Lai, Derek F. Wong, Lidia S. ChaoACL 2020 · 97 citations
- Self-supervised Masked Graph Autoencoder via Structure-aware CurriculumHaoyang Li, Xin Wang, Zeyang Zhang, Zongyuan Wu et al.ICML 2025
- Exploiting Structured Knowledge in Text via Graph-Guided Representation LearningTao Shen, Yi Mao, Pengcheng He, Guodong Long et al.EMNLP 2020 · 60 citations
- PMI-Masking: Principled masking of correlated spansYoav Levine, Barak Lenz, Opher Lieber, Omri Abend et al.ICLR 2021 · 83 citations
- Pre-Training Curriculum for Multi-Token Prediction in Language ModelsAnsar Aynetdinov, Alan AkbikACL 2025
