Single Teacher, Multiple Perspectives: Teacher Knowledge Augmentation for Enhanced Knowledge Distillation
Md. Imtiaz Hossain, Sharmen Akhter, Choong Seon Hong, Eui-Nam Huh
摘要
Do diverse perspectives help students learn better? Multi-teacher KD, which is a more effective technique than traditional single-teacher methods, supervises the student from different perspectives (i.e., teacher). While effective, multi-teacher, teacher ensemble, or teaching assistant-based approaches are computationally expensive and resourceintensive, as they require training multiple teacher networks. These concerns raise a question: can we supervise the student with diverse perspectives using only a single teacher? We, as the pioneer, demonstrate TeKAP, a novel teacher knowledge augmentation technique that generates multiple synthetic teacher knowledge by perturbing the knowledge of a single pretrained teacher i.e., Teacher Knowledge Augmentation via Perturbation, at both the feature and logit levels. These multiple augmented teachers simulate an ensemble of models together. The student model is trained on both the actual and augmented teacher knowledge, benefiting from the diversity of an ensemble without the need to train multiple teachers. TeKAP significantly reduces training time and computational resources, making it feasible for large-scale applications and easily manageable. Experimental results demonstrate that our proposed method helps existing state-of-the-art KD techniques achieve better performance, highlighting its potential as a cost-effective alternative. Source code can be found at: https://github.com/mdimtiazh/TeKAP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Local Dense Logit Relations for Enhanced Knowledge DistillationLiuchi Xu, Kang Liu, Jinshuai Liu, Lu Wang 等ICCV 2025 · 被引用 10 次
- Evidential Knowledge DistillationLiangyu Xiang, Junyu Gao, Changsheng XuICCV 2025 · 被引用 6 次
- Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular DiversitySeonghoon Yu, Dongjun Nam, Dina Katabi, Jeany SonNeurIPS 2025 · 被引用 4 次
- DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation TrainerHaiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju RenNeurIPS 2025 · 被引用 3 次
- Heterogeneous Complementary DistillationLiuchi Xu, Hao Zheng, Lu Wang, Lisheng Xu 等AAAI 2026
它引用的顶会 Paper12
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi 等NeurIPS 2021 · 被引用 318 次
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang 等CVPR 2024 · 被引用 183 次
- Densely Guided Knowledge Distillation using Multiple Teacher AssistantsWonchul Son, Jaemin Na, Junyong Choi, Wonjun HwangICCV 2021 · 被引用 158 次
相关 Paper
- Improving Ensemble Distillation With Weight Averaging and Diversifying PerturbationGiung Nam, Hyungi Lee, Byeongho Heo, Juho LeeICML 2022 · 被引用 10 次
- Class Incremental Learning with Multi-Teacher DistillationHaitao Wen, Lili Pan, Yu Dai, Heqian Qiu 等CVPR 2024 · 被引用 22 次
- How to Trade Off the Quantity and Capacity of Teacher Ensemble: Learning Categorical Distribution to Stochastically Employ a Teacher for DistillationZixiang Ding, Guoqing Jiang, Shuai Zhang, Lin Guo 等AAAI 2024 · 被引用 4 次
- Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion AugmentationMuquan Li, Dongyang Zhang, Tao He, Xiurui Xie 等ACM MM 2024 · 被引用 5 次
- Semi-supervised Learning with Multi-Head Co-TrainingMingcai Chen, Yuntao Du, Yi Zhang, Shuwei Qian 等AAAI 2022 · 被引用 42 次
