Class Attention Transfer Based Knowledge Distillation
Ziyao Guo, Haonan Yan, Hui Li, Xiaodong Lin
Abstract
Previous knowledge distillation methods have shown their impressive performance on model compression tasks, however, it is hard to explain how the knowledge they transferred helps to improve the performance of the student network. In this work, we focus on proposing a knowledge distillation method that has both high interpretability and competitive performance. We first revisit the structure of mainstream CNN models and reveal that possessing the capacity of identifying class discriminative regions of input is critical for CNN to perform classification. Furthermore, we demonstrate that this capacity can be obtained and enhanced by transferring class activation maps. Based on our findings, we propose class attention transfer based knowledge distillation (CAT-KD). Different from previous KD methods, we explore and present several properties of the knowledge transferred by our method, which not only improve the interpretability of CAT-KD but also contribute to a better understanding of CNN. While having high interpretability, CAT-KD achieves state-of-the-art performance on multiple benchmarks. Code is available at: https: //github.com/GzyAftermath/CAT-KD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2dc6d554-0f56-4441-aa02-5e96841565cfCited by top-tier papers31
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang et al.CVPR 2024 · 183 citations
- Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge DistillationJiaming Lv, Haoyuan Yang, Peihua LiNeurIPS 2024 · 59 citations
- Knowledge Distillation with Refined LogitsWujie Sun, Defang Chen, Siwei Lyu, Genlang Chen et al.ICCV 2025 · 13 citations
- Do Topological Characteristics Help in Knowledge Distillation?Jungeun Kim, Junwon You, Dongjin Lee, Ha Young Kim et al.ICML 2024 · 11 citations
- Cross-View Consistency Regularisation for Knowledge DistillationWeijia Zhang, Dongnan Liu, Weidong Cai, Chao MaACM MM 2024 · 11 citations
Builds on6
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park et al.ICCV 2019 · 727 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
- Towards Learning Spatially Discriminative Feature RepresentationsChaofei Wang, Jiayu Xiao, Yizeng Han, Qisen Yang et al.ICCV 2021 · 23 citations
Related papers
- On the Impact of Knowledge Distillation for Model InterpretabilityHyeongrok Han, Siwon Kim, Hyun-Soo Choi, Sungroh YoonICML 2023 · 13 citations
- Hierarchical Knowledge Squeezed Adversarial Network CompressionPeng Li, Chang Shu, Yuan Xie, Yan Qu et al.AAAI 2020 · 6 citations
- A Knowledge Distillation-Based Approach to Enhance Transparency of Classifier ModelsYuchen Jiang, Xinyuan Zhao, Yihang Wu, Ahmad ChaddadAAAI 2025 · 5 citations
- Multi-Knowledge Aggregation and Transfer for Semantic SegmentationYuang Liu, Wei Zhang, Jun WangAAAI 2022 · 11 citations
- Generative Model-Based Feature Knowledge Distillation for Action RecognitionGuiqin Wang, Peng Zhao, Yanjiang Shi, Cong Zhao et al.AAAI 2024 · 9 citations
