Supervision Complexity and its Role in Knowledge Distillation
Hrayr Harutyunyan, Ankit Singh Rawat, Aditya Krishna Menon, Seungyeon Kim, Sanjiv Kumar
Abstract
Despite the popularity and efficacy of knowledge distillation, there is limited understanding of why it helps. In order to study the generalization behavior of a distilled student, we propose a new theoretical framework that leverages supervision complexity: a measure of alignment between teacher-provided supervision and the student's neural tangent kernel. The framework highlights a delicate interplay among the teacher's accuracy, the student's margin with respect to the teacher predictions, and the complexity of the teacher predictions. Specifically, it provides a rigorous justification for the utility of various techniques that are prevalent in the context of distillation, such as early stopping and temperature scaling. Our analysis further suggests the use of online distillation, where a student receives increasingly more complex supervision from teachers in different stages of their training. We demonstrate efficacy of online distillation and validate the theoretical findings on a range of image classification benchmarks and model architectures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- On student-teacher deviations in distillation: does it pay to disobey?Vaishnavh Nagarajan, Aditya Krishna Menon, Srinadh Bhojanapalli, Hossein Mobahi et al.NeurIPS 2023 · 25 citations
- Cluster-aware Semi-supervised Learning: Relational Knowledge Distillation Provably Learns ClusteringYijun Dong, Kevin Miller, Qi Lei, Rachel A. WardNeurIPS 2023 · 9 citations
- MER-Inspector: Assessing Model Extraction Risks from An Attack-Agnostic PerspectiveXinwei Zhang, Haibo Hu, Qingqing Ye, Li Bai et al.WWW 2025 · 5 citations
- Towards the Fundamental Limits of Knowledge Transfer over Finite DomainsQingyue Zhao, Banghua ZhuICLR 2024 · 5 citations
- In Good GRACES: Principled Teacher Selection for Knowledge DistillationAbhishek Panigrahi, Bingbin Liu, Sadhika Malladi, Sham M. Kakade et al.ICLR 2026 · 5 citations
Builds on22
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
Related papers
- Knowledge Distillation in Wide Neural Networks: Risk Bound, Data Efficiency and Imperfect TeacherGuangda Ji, Zhanxing ZhuNeurIPS 2020 · 56 citations
- Distillation Scaling LawsDan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram et al.ICML 2025
- Neural Collapse Inspired Knowledge DistillationShuoxi Zhang, Zijian Song, Kun HeAAAI 2025 · 2 citations
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang et al.CVPR 2024 · 183 citations
- Exploring the Knowledge Transferred by Response-Based Teacher-Student DistillationLiangchen Song, Xuan Gong, Helong Zhou, Jiajie Chen et al.ACM MM 2023 · 14 citations
