Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient Space
Shangchen Du, Shan You, Xiaojie Li, Jianlong Wu, Fei Wang, Chen Qian, Changshui Zhang
Abstract
Distilling knowledge from an ensemble of teacher models is expected to have a more promising performance than that from a single one. Current methods mainly adopt a vanilla average rule, i.e., to simply take the average of all teacher losses for training the student network. However, this approach treats teachers equally and ignores the diversity among them. When conflicts or competitions exist among teachers, which is common, the inner compromise might hurt the distillation performance. In this paper, we examine the diversity of teacher models in the gradient space and regard the ensemble knowledge distillation as a multi-objective optimization problem so that we can determine a better optimization direction for the training of student network. Besides, we also introduce a tolerance parameter to accommodate disagreement among teachers. In this way, our method can be seen as a dynamic weighting method for each teacher in the ensemble. Extensive experiments validate the effectiveness of our method for both logits-based and feature-based cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8581a745-4b4a-41e2-a012-e7ae9572e35eCited by top-tier papers29
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian et al.NeurIPS 2022 · 477 citations
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang et al.CVPR 2022 · 213 citations
- LightTS: Lightweight Time Series Classification with Adaptive Ensemble DistillationDavid Campos, Miao Zhang, Bin Yang, Tung Kieu et al.SIGMOD 2023 · 105 citations
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 103 citations
- Student Customized Knowledge Distillation: Bridging the Gap Between Student and TeacherYichen Zhu, Yi WangICCV 2021 · 95 citations
Builds on8
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng et al.AAAI 2020 · 354 citations
- Deep Comprehensive Correlation Mining for Image ClusteringJianlong Wu, Keyu Long, Fei Wang, Chen Qian et al.ICCV 2019 · 191 citations
- ISTA-NAS: Efficient and Consistent Neural Architecture Search by Sparse CodingYibo Yang, Hongyang Li, Shan You, Fei Wang et al.NeurIPS 2020 · 66 citations
- Reborn Filters: Pruning Convolutional Neural Networks with Limited DataYehui Tang, Shan You, Chang Xu, Jin Han et al.AAAI 2020 · 33 citations
Related papers
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 10 citations
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon et al.ICML 2025
- Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular DiversitySeonghoon Yu, Dongjun Nam, Dina Katabi, Jeany SonNeurIPS 2025 · 4 citations
- How to Trade Off the Quantity and Capacity of Teacher Ensemble: Learning Categorical Distribution to Stochastically Employ a Teacher for DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo et al.AAAI 2024 · 4 citations
- DTO-KD: Dynamic Trade-off Optimization for Effective Knowledge DistillationZeeshan Hayder, Ali Cheraghian, Lars Petersson, Mehrtash Harandi et al.ICLR 2026
