Supervision Complexity and its Role in Knowledge Distillation
Hrayr Harutyunyan, Ankit Singh Rawat, Aditya Krishna Menon, Seungyeon Kim, Sanjiv Kumar
摘要
Despite the popularity and efficacy of knowledge distillation, there is limited understanding of why it helps. In order to study the generalization behavior of a distilled student, we propose a new theoretical framework that leverages supervision complexity: a measure of alignment between teacher-provided supervision and the student's neural tangent kernel. The framework highlights a delicate interplay among the teacher's accuracy, the student's margin with respect to the teacher predictions, and the complexity of the teacher predictions. Specifically, it provides a rigorous justification for the utility of various techniques that are prevalent in the context of distillation, such as early stopping and temperature scaling. Our analysis further suggests the use of online distillation, where a student receives increasingly more complex supervision from teachers in different stages of their training. We demonstrate efficacy of online distillation and validate the theoretical findings on a range of image classification benchmarks and model architectures.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- On student-teacher deviations in distillation: does it pay to disobey?Vaishnavh Nagarajan, Aditya Krishna Menon, Srinadh Bhojanapalli, Hossein Mobahi 等NeurIPS 2023 · 被引用 25 次
- Cluster-aware Semi-supervised Learning: Relational Knowledge Distillation Provably Learns ClusteringYijun Dong, Kevin Miller, Qi Lei, Rachel A. WardNeurIPS 2023 · 被引用 9 次
- MER-Inspector: Assessing Model Extraction Risks from An Attack-Agnostic PerspectiveXinwei Zhang, Haibo Hu, Qingqing Ye, Li Bai 等WWW 2025 · 被引用 5 次
- Towards the Fundamental Limits of Knowledge Transfer over Finite DomainsQingyue Zhao, Banghua ZhuICLR 2024 · 被引用 5 次
- In Good GRACES: Principled Teacher Selection for Knowledge DistillationAbhishek Panigrahi, Bingbin Liu, Sadhika Malladi, Sham M. Kakade 等ICLR 2026 · 被引用 5 次
它引用的顶会 Paper22
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
相关 Paper
- Knowledge Distillation in Wide Neural Networks: Risk Bound, Data Efficiency and Imperfect TeacherGuangda Ji, Zhanxing ZhuNeurIPS 2020 · 被引用 56 次
- Distillation Scaling LawsDan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram 等ICML 2025
- Neural Collapse Inspired Knowledge DistillationShuoxi Zhang, Zijian Song, Kun HeAAAI 2025 · 被引用 2 次
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang 等CVPR 2024 · 被引用 183 次
- Exploring the Knowledge Transferred by Response-Based Teacher-Student DistillationLiangchen Song, Xuan Gong, Helong Zhou, Jiajie Chen 等ACM MM 2023 · 被引用 14 次
