Reinforced Multi-Teacher Selection for Knowledge Distillation
Fei Yuan, Linjun Shou, Jian Pei, Wutao Lin, Ming Gong, Yan Fu, Daxin Jiang
Abstract
In natural language processing (NLP) tasks, slow inference speed and huge footprints in GPU usage remain the bottleneck of applying pre-trained deep models in production. As a popular method for model compression, knowledge distillation transfers knowledge from one or multiple large (teacher) models to a small (student) model. When multiple teacher models are available in distillation, the state-of-the-art methods assign a fixed weight to a teacher model in the whole distillation. Furthermore, most of the existing methods allocate an equal weight to every teacher model. In this paper, we observe that, due to the complexity of training examples and the differences in student model capability, learning differentially from teacher models can lead to better performance of student models distilled. We systematically develop a reinforced method to dynamically assign weights to teacher models for different training instances and optimize the performance of student model. Our extensive experimental results on several NLP tasks clearly verify the feasibility and effectiveness of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 994c9835-9e75-40ad-84b8-283e44147efaCited by top-tier papers22
- LightTS: Lightweight Time Series Classification with Adaptive Ensemble DistillationDavid Campos, Miao Zhang, Bin Yang, Tung Kieu et al.SIGMOD 2023 · 105 citations
- Reinforcement Learning Based Dynamic Model Combination for Time Series ForecastingYuwei Fu, Di Wu, Benoit BouletAAAI 2022 · 66 citations
- Distillation from Heterogeneous Models for Top-K RecommendationSeongKu Kang, Wonbin Kweon, Dongha Lee, Jianxun Lian et al.WWW 2023 · 35 citations
- AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into OneMike Ranzinger, Greg Heinrich, Jan Kautz, Pavlo MolchanovCVPR 2024 · 31 citations
- FreeKD: Free-direction Knowledge Distillation for Graph Neural NetworksKaituo Feng, Changsheng Li, Ye Yuan, Guoren WangKDD 2022 · 28 citations
Builds on1
Related papers
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
- Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual RecognitionChuanguang Yang, Xinqiang Yu, Han Yang, Zhulin An et al.AAAI 2025 · 26 citations
- Dynamic Knowledge Distillation for Pre-trained Language ModelsLei Li, Yankai Lin, Shuhuai Ren, Peng Li et al.EMNLP 2021 · 32 citations
- Weight Distillation: Transferring the Knowledge in Neural Network ParametersYe Lin, Yanyang Li, Ziyang Wang, Bei Li et al.ACL 2021
- Masking Teacher and Reinforcing Student for Distilling Vision-Language ModelsByung-Kwan Lee, Yu-Chiang Frank Wang, Ryo HachiumaCVPR 2026 · 7 citations
