Dynamic Knowledge Distillation for Pre-trained Language Models
Lei Li, Yankai Lin, Shuhuai Ren, Peng Li, Jie Zhou, Xu Sun
Abstract
Knowledge distillation (KD) has been proved effective for compressing large-scale pretrained language models. However, existing methods conduct KD statically, e.g., the student model aligns its output distribution to that of a selected teacher model on the pre-defined training dataset. In this paper, we explore whether a dynamic knowledge distillation that empowers the student to adjust the learning procedure according to its competency, regarding the student performance and learning efficiency. We explore the dynamical adjustments on three aspects: teacher model adoption, data selection, and KD objective adaptation. Experimental results show that (1) proper selection of teacher model can boost the performance of student model; (2) conducting KD with 10% informative instances achieves comparable performance while greatly accelerates the training; (3) the student performance can be boosted by adjusting the supervision contribution of different alignment objective. We find dynamic knowledge distillation is promising and provide discussions on potential future directions towards more efficient KD methods. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8a4a445-7acd-40df-ac0a-199b57f68954Cited by top-tier papers9
- D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language ModelsYikun Jiang, Huanyu Wang, Lei Xie, Hanbin Zhao et al.NeurIPS 2024 · 39 citations
- Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language ModelsLei Li, Yuqi Wang, Runxin Xu, Peiyi Wang et al.ACL 2024 · 16 citations
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time ScalingSachin Goyal, David Lopez-Paz, Kartik AhujaICLR 2026 · 11 citations
- HomoDistil: Homotopic Task-Agnostic Distillation of Pre-trained TransformersChen Liang, Haoming Jiang, Zheng Li, Xianfeng Tang et al.ICLR 2023 · 9 citations
- Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor NetworkJunho Kim, Jun-Hyung Park, Mingyu Lee, Wing-Lam Mok et al.EMNLP 2022 · 4 citations
Builds on8
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot et al.ICLR 2020 · 244 citations
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei et al.EMNLP 2020 · 168 citations
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou et al.ICLR 2021 · 90 citations
Related papers
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang et al.NeurIPS 2024 · 50 citations
- Reinforced Multi-Teacher Selection for Knowledge DistillationFei Yuan, Linjun Shou, Jian Pei, Wutao Lin et al.AAAI 2021 · 155 citations
- Revisiting Knowledge Distillation for Autoregressive Language ModelsQihuang Zhong, Liang Ding, Li Shen, Juhua Liu et al.ACL 2024
- Can Students Beyond the Teacher? Distilling Knowledge from Teacher's BiasJianhua Zhang, Yi Gao, Ruyu Liu, Xu Cheng et al.AAAI 2025 · 2 citations
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMsHaiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju RenCVPR 2026 · 18 citations
