Dynamic Knowledge Distillation for Pre-trained Language Models
Lei Li, Yankai Lin, Shuhuai Ren, Peng Li, Jie Zhou, Xu Sun
摘要
Knowledge distillation (KD) has been proved effective for compressing large-scale pretrained language models. However, existing methods conduct KD statically, e.g., the student model aligns its output distribution to that of a selected teacher model on the pre-defined training dataset. In this paper, we explore whether a dynamic knowledge distillation that empowers the student to adjust the learning procedure according to its competency, regarding the student performance and learning efficiency. We explore the dynamical adjustments on three aspects: teacher model adoption, data selection, and KD objective adaptation. Experimental results show that (1) proper selection of teacher model can boost the performance of student model; (2) conducting KD with 10% informative instances achieves comparable performance while greatly accelerates the training; (3) the student performance can be boosted by adjusting the supervision contribution of different alignment objective. We find dynamic knowledge distillation is promising and provide discussions on potential future directions towards more efficient KD methods. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language ModelsYikun Jiang, Huanyu Wang, Lei Xie, Hanbin Zhao 等NeurIPS 2024 · 被引用 39 次
- Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language ModelsLei Li, Yuqi Wang, Runxin Xu, Peiyi Wang 等ACL 2024 · 被引用 16 次
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time ScalingSachin Goyal, David Lopez-Paz, Kartik AhujaICLR 2026 · 被引用 11 次
- HomoDistil: Homotopic Task-Agnostic Distillation of Pre-trained TransformersChen Liang, Haoming Jiang, Zheng Li, Xianfeng Tang 等ICLR 2023 · 被引用 9 次
- Tutoring Helps Students Learn Better: Improving Knowledge Distillation for BERT with Tutor NetworkJunho Kim, Jun-Hyung Park, Mingyu Lee, Wing-Lam Mok 等EMNLP 2022 · 被引用 4 次
它引用的顶会 Paper8
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Thieves on Sesame Street! Model Extraction of BERT-based APIsKalpesh Krishna, Gaurav Singh Tomar, Ankur P. Parikh, Nicolas Papernot 等ICLR 2020 · 被引用 244 次
- BERT-of-Theseus: Compressing BERT by Progressive Module ReplacingCanwen Xu, Wangchunshu Zhou, Tao Ge, Furu Wei 等EMNLP 2020 · 被引用 168 次
- MixKD: Towards Efficient Distillation of Large-scale Language ModelsKevin J. Liang, Weituo Hao, Dinghan Shen, Yufan Zhou 等ICLR 2021 · 被引用 90 次
相关 Paper
- DDK: Distilling Domain Knowledge for Efficient Large Language ModelsJiaheng Liu, Chenchen Zhang, Jinyang Guo, Yuanxing Zhang 等NeurIPS 2024 · 被引用 50 次
- Reinforced Multi-Teacher Selection for Knowledge DistillationFei Yuan, Linjun Shou, Jian Pei, Wutao Lin 等AAAI 2021 · 被引用 155 次
- Revisiting Knowledge Distillation for Autoregressive Language ModelsQihuang Zhong, Liang Ding, Li Shen, Juhua Liu 等ACL 2024
- Can Students Beyond the Teacher? Distilling Knowledge from Teacher's BiasJianhua Zhang, Yi Gao, Ruyu Liu, Xu Cheng 等AAAI 2025 · 被引用 2 次
- SelecTKD: Selective Token-Weighted Knowledge Distillation for LLMsHaiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju RenCVPR 2026 · 被引用 18 次
