Reinforced Multi-Teacher Selection for Knowledge Distillation
Fei Yuan, Linjun Shou, Jian Pei, Wutao Lin, Ming Gong, Yan Fu, Daxin Jiang
摘要
In natural language processing (NLP) tasks, slow inference speed and huge footprints in GPU usage remain the bottleneck of applying pre-trained deep models in production. As a popular method for model compression, knowledge distillation transfers knowledge from one or multiple large (teacher) models to a small (student) model. When multiple teacher models are available in distillation, the state-of-the-art methods assign a fixed weight to a teacher model in the whole distillation. Furthermore, most of the existing methods allocate an equal weight to every teacher model. In this paper, we observe that, due to the complexity of training examples and the differences in student model capability, learning differentially from teacher models can lead to better performance of student models distilled. We systematically develop a reinforced method to dynamically assign weights to teacher models for different training instances and optimize the performance of student model. Our extensive experimental results on several NLP tasks clearly verify the feasibility and effectiveness of our approach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- LightTS: Lightweight Time Series Classification with Adaptive Ensemble DistillationDavid Campos, Miao Zhang, Bin Yang, Tung Kieu 等SIGMOD 2023 · 被引用 105 次
- Reinforcement Learning Based Dynamic Model Combination for Time Series ForecastingYuwei Fu, Di Wu, Benoit BouletAAAI 2022 · 被引用 66 次
- Distillation from Heterogeneous Models for Top-K RecommendationSeongKu Kang, Wonbin Kweon, Dongha Lee, Jianxun Lian 等WWW 2023 · 被引用 35 次
- AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into OneMike Ranzinger, Greg Heinrich, Jan Kautz, Pavlo MolchanovCVPR 2024 · 被引用 31 次
- FreeKD: Free-direction Knowledge Distillation for Graph Neural NetworksKaituo Feng, Changsheng Li, Ye Yuan, Guoren WangKDD 2022 · 被引用 28 次
它引用的顶会 Paper1
相关 Paper
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang 等ACL 2021
- Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual RecognitionChuanguang Yang, Xinqiang Yu, Han Yang, Zhulin An 等AAAI 2025 · 被引用 26 次
- Dynamic Knowledge Distillation for Pre-trained Language ModelsLei Li, Yankai Lin, Shuhuai Ren, Peng Li 等EMNLP 2021 · 被引用 32 次
- Weight Distillation: Transferring the Knowledge in Neural Network ParametersYe Lin, Yanyang Li, Ziyang Wang, Bei Li 等ACL 2021
- Masking Teacher and Reinforcing Student for Distilling Vision-Language ModelsByung-Kwan Lee, Yu-Chiang Frank Wang, Ryo HachiumaCVPR 2026 · 被引用 7 次
