Understanding the Role of the Projector in Knowledge Distillation
Roy Miles, Krystian Mikolajczyk
Abstract
In this paper we revisit the efficacy of knowledge distillation as a function matching and metric learning problem. In doing so we verify three important design decisions, namely the normalisation, soft maximum function, and projection layers as key ingredients. We theoretically show that the projector implicitly encodes information on past examples, enabling relational gradients for the student. We then show that the normalisation of representations is tightly coupled with the training dynamics of this projector, which can have a large impact on the students performance. Finally, we show that a simple soft maximum function can be used to address any significant capacity gap problems. Experimental results on various benchmark datasets demonstrate that using these insights can lead to superior or comparable performance to state-of-the-art knowledge distillation techniques, despite being much more computationally efficient. In particular, we obtain these results across image classification (CIFAR100 and ImageNet), object detection (COCO2017), and on more difficult distillation objectives, such as training data efficient transformers, whereby we attain a 77.2% top-1 accuracy with DeiT-Ti on ImageNet. Code and models are publicly available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 717b16c8-b68f-4114-a3c7-3ca66e5b6689Cited by top-tier papers11
- Relational Diffusion Distillation for Efficient Image GenerationWeilun Feng, Chuanguang Yang, Zhulin An, Libo Huang et al.ACM MM 2024 · 11 citations
- Local Dense Logit Relations for Enhanced Knowledge DistillationLiuchi Xu, Kang Liu, Jinshuai Liu, Lu Wang et al.ICCV 2025 · 10 citations
- : Improving Knowledge Distillation Using Orthogonal ProjectionsRoy Miles, Ismail Elezi, Jiankang DengCVPR 2024 · 9 citations
- Debiased Distillation for Consistency RegularizationLu Wang, Liuchi Xu, Xiong Yang, Zhenhua Huang et al.AAAI 2025 · 6 citations
- Hybrid Data-Free Knowledge DistillationJialiang Tang, Shuo Chen, Chen GongAAAI 2025 · 2 citations
Builds on17
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 4,453 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- DTO-KD: Dynamic Trade-off Optimization for Effective Knowledge DistillationZeeshan Hayder, Ali Cheraghian, Lars Petersson, Mehrtash Harandi et al.ICLR 2026
- Knowledge Distillation with the Reused Teacher ClassifierDefang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang et al.CVPR 2022 · 213 citations
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang et al.CVPR 2024 · 183 citations
- Improved Feature Distillation via Projector EnsembleYudong Chen, Sen Wang, Jiajun Liu, Xuwei Xu et al.NeurIPS 2022 · 73 citations
- Better Teacher Better Student: Dynamic Prior Knowledge for Knowledge DistillationMartin Zong, Zengyu Qiu, Xinzhu Ma, Kunlin Yang et al.ICLR 2023 · 19 citations
