Show, Attend and Distill: Knowledge Distillation via Attention-based Feature Matching
Mingi Ji, Byeongho Heo, Sungrae Park
Abstract
Knowledge distillation extracts general knowledge from a pretrained teacher network and provides guidance to a target student network. Most studies manually tie intermediate features of the teacher and student, and transfer knowledge through predefined links. However, manual selection often constructs ineffective links that limit the improvement from the distillation. There has been an attempt to address the problem, but it is still challenging to identify effective links under practical scenarios. In this paper, we introduce an effective and efficient feature distillation method utilizing all the feature levels of the teacher without manually selecting the links. Specifically, our method utilizes an attention-based meta-network that learns relative similarities between features, and applies identified similarities to control distillation intensities of all possible pairs. As a result, our method determines competent links more efficiently than the previous approach and provides better performance on model compression and transfer learning tasks. Further qualitative analyses and ablative studies describe how our method contributes to better distillation. The implementation code is available at github.com/clovaai/attention-feature-distillation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers35
- Exploring Inter-Channel Correlation for Diversity-preserved Knowledge DistillationLi Liu, Qingle Huang, Sihao Lin, Hongwei Xie et al.ICCV 2021 · 132 citations
- Knowledge Distillation via the Target-aware TransformerSihao Lin, Hongwei Xie, Bing Wang, Kaicheng Yu et al.CVPR 2022 · 126 citations
- Student Customized Knowledge Distillation: Bridging the Gap Between Student and TeacherYichen Zhu, Yi WangICCV 2021 · 95 citations
- Improved Feature Distillation via Projector EnsembleYudong Chen, Sen Wang, Jiajun Liu, Xuwei Xu et al.NeurIPS 2022 · 73 citations
- Improving Adversarial Robustness via Information Bottleneck DistillationHuafeng Kuang, Hong Liu, Yongjian Wu, Shin'ichi Satoh et al.NeurIPS 2023 · 27 citations
Builds on2
Related papers
- Cross-Layer Distillation with Semantic CalibrationDefang Chen, Jian-Ping Mei, Yuan Zhang, Can Wang et al.AAAI 2021 · 368 citations
- Towards Efficient Pre-Trained Language Model via Feature Correlation DistillationKun Huang, Xin Guo, Meng WangNeurIPS 2023 · 8 citations
- Meta-KD: A Meta Knowledge Distillation Framework for Language Model Compression across DomainsHaojie Pan, Chengyu Wang, Minghui Qiu, Yichang Zhang et al.ACL 2021
- Distilling Knowledge via Knowledge ReviewPengguang Chen, Shu Liu, Hengshuang Zhao, Jiaya JiaCVPR 2021
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li et al.ACL 2025
