A statistical perspective on distillation
Aditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim, Sanjiv Kumar
Abstract
Knowledge distillation is a technique for improving a "student" model by replacing its one-hot training labels with a label distribution obtained from a "teacher" model. Despite its broad success, several basic questions -e.g., Why does distillation help? Why do more accurate teachers not necessarily distill better? -have received limited formal study. In this paper, we present a statistical perspective on distillation which sheds light on these questions. Our core observation is that a "Bayes teacher" providing the true classprobabilities can lower the variance of the student objective, and thus improve performance. We then establish a bias-variance tradeoff that quantifies the utility of teachers that approximate the Bayes class-probabilities. This provides a formal criterion as to what constitutes a "good" teacher, namely, the quality of its probability estimates. Finally, we illustrate how our statistical perspective facilitates novel applications of distillation to bipartite ranking and multiclass retrieval.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7cdf99a-db92-4991-a96d-d3fa90c701f7Cited by top-tier papers42
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 102 citations
- Better Uncertainty Calibration via Proper Scores for Classification and BeyondSebastian G. Gruber, Florian BuettnerNeurIPS 2022 · 88 citations
- What Makes a "Good" Data Augmentation in Knowledge Distillation - A Statistical PerspectiveHuan Wang, Suhas Lohit, Michael J. Jones, Yun FuNeurIPS 2022 · 62 citations
- What Knowledge Gets Distilled in Knowledge Distillation?Utkarsh Ojha, Yuheng Li, Anirudh Sundara Rajan, Yingyu Liang et al.NeurIPS 2023 · 58 citations
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li et al.NeurIPS 2022 · 46 citations
Builds on11
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 411 citations
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 298 citations
- Rethinking Bias-Variance Trade-off for Generalization of Neural NetworksZitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt et al.ICML 2020 · 219 citations
Related papers
- SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and GuidelinesItai Morad, Nir Shlezinger, Yonina C. EldarICLR 2026 · 1 citation
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi et al.NeurIPS 2021 · 318 citations
- PLD: A Choice-Theoretic List-Wise Knowledge DistillationEjafa Bassam, Dawei Zhu, Kaigui BianNeurIPS 2025
- Knowledge Distillation with Perturbed Loss: From a Vanilla Teacher to a Proxy TeacherRongzhi Zhang, Jiaming Shen, Tianqi Liu, Jialu Liu et al.KDD 2024 · 3 citations
- On student-teacher deviations in distillation: does it pay to disobey?Vaishnavh Nagarajan, Aditya Krishna Menon, Srinadh Bhojanapalli, Hossein Mobahi et al.NeurIPS 2023 · 25 citations
