MiniPLM: Knowledge Distillation for Pre-training Language Models
Yuxian Gu, Hao Zhou, Fandong Meng, Jie Zhou, Minlie Huang
Abstract
Knowledge distillation (KD) is widely used to train small, high-performing student language models (LMs) using large teacher LMs. While effective in fine-tuning, KD during pre-training faces challenges in efficiency, flexibility, and effectiveness. Existing methods either incur high computational costs due to online teacher inference, require tokenization matching between teacher and student LMs, or risk losing the difficulty and diversity of the teacher-generated training data. To address these issues, we propose MINIPLM, a KD framework for pre-training LMs by refining the training data distribution with the teacher's knowledge. For efficiency, MINIPLM performs offline teacher LM inference, allowing KD for multiple student LMs without adding training-time costs. For flexibility, MINIPLM operates solely on the training corpus, enabling KD across model families. For effectiveness, MINIPLM leverages the differences between large and small LMs to enhance the difficulty and diversity of the training data, helping student LMs acquire versatile and sophisticated knowledge. Extensive experiments demonstrate that MINIPLM boosts the student LMs' performance on 9 widely used downstream tasks, improves the language modeling capabilities, and reduces pre-training computation. The benefit of MINIPLM extends to large pre-training scales, evidenced by the extrapolation of the scaling curves. Further analysis reveals that MINIPLM supports KD across model families and enhances the utilization of pre-training data. Our model, code, and data are available at https://github.com/thu-coai/MiniPLM . 2.2x (a) Computation Scaling (1.8B → 500M) * Contribution during an internship at Tencent Inc.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- LLM Pretraining with Continuous ConceptsJihoon Tack, Jack Lanchantin, Jane Dwivedi-Yu, Andrew Cohen et al.ICLR 2026 · 30 citations
- A Token is Worth over 1, 000 Tokens: Efficient Knowledge Distillation through Low-Rank CloneJitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao et al.NeurIPS 2025 · 17 citations
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time ScalingSachin Goyal, David Lopez-Paz, Kartik AhujaICLR 2026 · 11 citations
- PASER: Post-Training Data Selection for Efficient Pruned Large Language Model RecoveryBowei He, Lihao Yin, Huiling Zhen, Xiaokun Zhang et al.ICLR 2026 · 5 citations
- STAT: Skill-Targeted Adaptive TrainingYinghui He, Abhishek Panigrahi, Yong Lin, Sanjeev AroraICLR 2026 · 3 citations
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du et al.NeurIPS 2023 · 457 citations
Related papers
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 95 citations
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
- Pre-training Distillation for Large Language Models: A Design Space ExplorationHao Peng, Xin Lv, Yushi Bai, Zijun Yao et al.ACL 2025
- An Empirical Study of Knowledge Distillation for Code Understanding TasksRuiqi Wang, Zezhou Yang, Cuiyun Gao, Xin Xia et al.ICSE 2026 · 1 citation
- DA-KD: Difficulty-Aware Knowledge Distillation for Efficient Large Language ModelsChangyi He, Yifu Ding, Jinyang Guo, Ruihao Gong et al.ICML 2025
