Avatar Knowledge Distillation: Self-ensemble Teacher Paradigm with Uncertainty
Yuan Zhang, Weihua Chen, Yichen Lu, Tao Huang, Xiuyu Sun, Jian Cao
Abstract
Knowledge distillation is an effective paradigm for boosting the performance of pocket-size model, especially when multiple teacher models are available, the student would break the upper limit again. However, it is not economical to train diverse teacher models for the disposable distillation. In this paper, we introduce a new concept dubbed Avatars for distillation, which are the inference ensemble models derived from the teacher. Concretely, (1) For each iteration of distillation training, various Avatars are generated by a perturbation transformation. We validate that Avatars own higher upper limit of working capacity and teaching ability, aiding the student model in learning diverse and receptive knowledge perspectives from the teacher model. (2) During the distillation, we propose an uncertainty-aware factor from the variance of statistical differences between the vanilla teacher and Avatars, to adjust Avatars' contribution on knowledge transfer adaptively. Avatar Knowledge Distillation (AKD) is fundamentally different from existing methods and refines with the innovative view of unequal training. Comprehensive experiments demonstrate the effectiveness of our Avatars mechanism, which polishes up the state-of-the-art distillation methods for dense prediction without more extra computational cost. The AKD brings at most 0.7 AP gains on COCO 2017 for Object Detection and 1.83 mIoU gains on Cityscapes for Semantic Segmentation, respectively. Code is
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bd8dcbaa-36a5-42a2-95b9-b39a35ca3cddCited by top-tier papers5
- Knowledge Diffusion for DistillationTao Huang, Yuan Zhang, Mingkai Zheng, Shan You et al.NeurIPS 2023 · 125 citations
- FreeKD: Knowledge Distillation via Semantic Frequency PromptYuan Zhang, Tao Huang, Jiaming Liu, Tao Jiang et al.CVPR 2024 · 26 citations
- Cloud-Device Collaborative Learning for Multimodal Large Language ModelsGuanqun Wang, Jiaming Liu, Chenxuan Li, Yuan Zhang et al.CVPR 2024 · 9 citations
- Evidential Knowledge DistillationLiangyu Xiang, Junyu Gao, Changsheng XuICCV 2025 · 6 citations
- Distilling Cross-Modal Knowledge via Feature DisentanglementJunhong Liu, Yuan Zhang, Tao Huang, Wenchao Xu et al.AAAI 2026 · 2 citations
Builds on18
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- RepPoints: Point Set Representation for Object DetectionZe Yang, Shaohui Liu, Han Hu, Liwei Wang et al.ICCV 2019 · 1,056 citations
- An Empirical Study of Spatial Attention Mechanisms in Deep NetworksXizhou Zhu, Dazhi Cheng, Zheng Zhang, Stephen Lin et al.ICCV 2019 · 522 citations
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian et al.NeurIPS 2022 · 477 citations
- Channel-wise Knowledge Distillation for Dense Prediction*Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan et al.ICCV 2021 · 432 citations
Related papers
- UniKD: Universal Knowledge Distillation for Mimicking Homogeneous or Heterogeneous Object DetectorsShanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu et al.ICCV 2023 · 7 citations
- Localization Distillation for Dense Object DetectionZhaohui Zheng, Rongguang Ye, Ping Wang, Dongwei Ren et al.CVPR 2022 · 177 citations
- ScaleKD: Distilling Scale-Aware Knowledge in Small Object DetectorYichen Zhu, Qiqi Zhou, Ning Liu, Zhiyuan Xu et al.CVPR 2023
- DTO-KD: Dynamic Trade-off Optimization for Effective Knowledge DistillationZeeshan Hayder, Ali Cheraghian, Lars Petersson, Mehrtash Harandi et al.ICLR 2026
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng et al.AAAI 2020 · 354 citations
