Pre-training Distillation for Large Language Models: A Design Space Exploration
Hao Peng, Xin Lv, Yushi Bai, Zijun Yao, Jiajie Zhang, Lei Hou, Juanzi Li
Abstract
Knowledge distillation (KD) aims to transfer knowledge from a large teacher model to a smaller student model. Previous work applying KD in the field of large language models (LLMs) typically focused on the post-training phase, where the student LLM learns directly from instructions and corresponding responses generated by the teacher model. In this paper, we extend KD to the pre-training phase of LLMs, named pre-training distillation (PD). We first conduct a preliminary experiment using GLM-4-9B as the teacher LLM to distill a 1.9B parameter student LLM, validating the effectiveness of PD. Considering the key impact factors of distillation, we systematically explore the design space of pre-training distillation across four aspects: logits processing, loss selection, scaling law, and offline or online logits. We conduct extensive experiments to explore the design space of pre-training distillation and find better configurations and interesting conclusions, such as larger student LLMs generally benefiting more from pre-training distillation, while a larger teacher LLM does not necessarily guarantee better results. We hope our exploration of the design space will inform future practices in pre-training distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62d781f3-c4d3-4893-80d2-e5ea1452a28cCited by top-tier papers7
- Entropy-Aware On-Policy Distillation of Language ModelsWoogyeol Jin, Taywon Min, Yongjin Yang, Dennis Wei et al.ICML 2026 · 91 citations
- In Good GRACES: Principled Teacher Selection for Knowledge DistillationAbhishek Panigrahi, Bingbin Liu, Sadhika Malladi, Sham M. Kakade et al.ICLR 2026 · 5 citations
- Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMsAnshumann, Mohd Abbas Zaidi, Akhil Kedia, Jinwoo Ahn et al.ACL 2025
- STAR: Similarity-guided Teacher-Assisted Refinement for Super-Tiny Function Calling ModelsJiliang Ni, Jiachen Pu, Zhongyi Yang, Jingfeng Luo et al.ICLR 2026
- Comprehensive and Efficient Distillation for Lightweight Sentiment Analysis ModelsGuangyu Xie, Yice Zhang, Jianzhu Bao, Qianlong Wang et al.EMNLP 2025
Builds on22
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
Related papers
- MiniPLM: Knowledge Distillation for Pre-training Language ModelsYuxian Gu, Hao Zhou, Fandong Meng, Jie Zhou et al.ICLR 2025
- An Empirical Study of Knowledge Distillation for Code Understanding TasksRuiqi Wang, Zezhou Yang, Cuiyun Gao, Xin Xia et al.ICSE 2026 · 1 citation
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang et al.NeurIPS 2024 · 50 citations
- Why Knowledge Distillation Works in Generative Models: A Minimal Working ExplanationSungmin Cha, Kyunghyun ChoNeurIPS 2025 · 9 citations
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
