MiniPLM: Knowledge Distillation for Pre-training Language Models
Yuxian Gu, Hao Zhou, Fandong Meng, Jie Zhou, Minlie Huang
摘要
Knowledge distillation (KD) is widely used to train small, high-performing student language models (LMs) using large teacher LMs. While effective in fine-tuning, KD during pre-training faces challenges in efficiency, flexibility, and effectiveness. Existing methods either incur high computational costs due to online teacher inference, require tokenization matching between teacher and student LMs, or risk losing the difficulty and diversity of the teacher-generated training data. To address these issues, we propose MINIPLM, a KD framework for pre-training LMs by refining the training data distribution with the teacher's knowledge. For efficiency, MINIPLM performs offline teacher LM inference, allowing KD for multiple student LMs without adding training-time costs. For flexibility, MINIPLM operates solely on the training corpus, enabling KD across model families. For effectiveness, MINIPLM leverages the differences between large and small LMs to enhance the difficulty and diversity of the training data, helping student LMs acquire versatile and sophisticated knowledge. Extensive experiments demonstrate that MINIPLM boosts the student LMs' performance on 9 widely used downstream tasks, improves the language modeling capabilities, and reduces pre-training computation. The benefit of MINIPLM extends to large pre-training scales, evidenced by the extrapolation of the scaling curves. Further analysis reveals that MINIPLM supports KD across model families and enhances the utilization of pre-training data. Our model, code, and data are available at https://github.com/thu-coai/MiniPLM . 2.2x (a) Computation Scaling (1.8B → 500M) * Contribution during an internship at Tencent Inc.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- LLM Pretraining with Continuous ConceptsJihoon Tack, Jack Lanchantin, Jane Dwivedi-Yu, Andrew Cohen 等ICLR 2026 · 被引用 30 次
- A Token is Worth over 1, 000 Tokens: Efficient Knowledge Distillation through Low-Rank CloneJitai Hao, Qiang Huang, Hao Liu, Xinyan Xiao 等NeurIPS 2025 · 被引用 17 次
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time ScalingSachin Goyal, David Lopez-Paz, Kartik AhujaICLR 2026 · 被引用 11 次
- PASER: Post-Training Data Selection for Efficient Pruned Large Language Model RecoveryBowei He, Lihao Yin, Huiling Zhen, Xiaokun Zhang 等ICLR 2026 · 被引用 5 次
- STAT: Skill-Targeted Adaptive TrainingYinghui He, Abhishek Panigrahi, Yong Lin, Sanjeev AroraICLR 2026 · 被引用 3 次
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng 等ICML 2020 · 被引用 1,388 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du 等NeurIPS 2023 · 被引用 457 次
相关 Paper
- MiniLLM: Knowledge Distillation of Large Language ModelsYuxian Gu, Li Dong, Furu Wei, Minlie HuangICLR 2024 · 被引用 95 次
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang 等ACL 2021
- Pre-training Distillation for Large Language Models: A Design Space ExplorationHao Peng, Xin Lv, Yushi Bai, Zijun Yao 等ACL 2025
- An Empirical Study of Knowledge Distillation for Code Understanding TasksRuiqi Wang, Zezhou Yang, Cuiyun Gao, Xin Xia 等ICSE 2026 · 被引用 1 次
- DA-KD: Difficulty-Aware Knowledge Distillation for Efficient Large Language ModelsChangyi He, Yifu Ding, Jinyang Guo, Ruihao Gong 等ICML 2025
