Agree to Disagree: Adaptive Ensemble Knowledge Distillation in Gradient Space
Shangchen Du, Shan You, Xiaojie Li, Jianlong Wu, Fei Wang, Chen Qian, Changshui Zhang
摘要
Distilling knowledge from an ensemble of teacher models is expected to have a more promising performance than that from a single one. Current methods mainly adopt a vanilla average rule, i.e., to simply take the average of all teacher losses for training the student network. However, this approach treats teachers equally and ignores the diversity among them. When conflicts or competitions exist among teachers, which is common, the inner compromise might hurt the distillation performance. In this paper, we examine the diversity of teacher models in the gradient space and regard the ensemble knowledge distillation as a multi-objective optimization problem so that we can determine a better optimization direction for the training of student network. Besides, we also introduce a tolerance parameter to accommodate disagreement among teachers. In this way, our method can be seen as a dynamic weighting method for each teacher in the ensemble. Extensive experiments validate the effectiveness of our method for both logits-based and feature-based cases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian 等NeurIPS 2022 · 被引用 477 次
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang 等CVPR 2022 · 被引用 213 次
- LightTS: Lightweight Time Series Classification with Adaptive Ensemble DistillationDavid Campos, Miao Zhang, Bin Yang, Tung Kieu 等SIGMOD 2023 · 被引用 105 次
- Shadow Knowledge Distillation: Bridging Offline and Online Knowledge TransferLujun Li, Zhe JinNeurIPS 2022 · 被引用 103 次
- Student Customized Knowledge Distillation: Bridging the Gap Between Student and TeacherYichen Zhu, Yi WangICCV 2021 · 被引用 95 次
它引用的顶会 Paper8
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou 等ICCV 2019 · 被引用 625 次
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng 等AAAI 2020 · 被引用 354 次
- Deep Comprehensive Correlation Mining for Image ClusteringJianlong Wu, Keyu Long, Fei Wang, Chen Qian 等ICCV 2019 · 被引用 191 次
- ISTA-NAS: Efficient and Consistent Neural Architecture Search by Sparse CodingYibo Yang, Hongyang Li, Shan You, Fei Wang 等NeurIPS 2020 · 被引用 66 次
- Reborn Filters: Pruning Convolutional Neural Networks with Limited DataYehui Tang, Shan You, Chang Xu, Jin Han 等AAAI 2020 · 被引用 33 次
相关 Paper
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 被引用 10 次
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon 等ICML 2025
- Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular DiversitySeonghoon Yu, Dongjun Nam, Dina Katabi, Jeany SonNeurIPS 2025 · 被引用 4 次
- How to Trade Off the Quantity and Capacity of Teacher Ensemble: Learning Categorical Distribution to Stochastically Employ a Teacher for DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo 等AAAI 2024 · 被引用 4 次
- DTO-KD: Dynamic Trade-off Optimization for Effective Knowledge DistillationZeeshan Hayder, Ali Cheraghian, Lars Petersson, Mehrtash Harandi 等ICLR 2026
