A statistical perspective on distillation
Aditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim, Sanjiv Kumar
摘要
Knowledge distillation is a technique for improving a "student" model by replacing its one-hot training labels with a label distribution obtained from a "teacher" model. Despite its broad success, several basic questions -e.g., Why does distillation help? Why do more accurate teachers not necessarily distill better? -have received limited formal study. In this paper, we present a statistical perspective on distillation which sheds light on these questions. Our core observation is that a "Bayes teacher" providing the true classprobabilities can lower the variance of the student objective, and thus improve performance. We then establish a bias-variance tradeoff that quantifies the utility of teachers that approximate the Bayes class-probabilities. This provides a formal criterion as to what constitutes a "good" teacher, namely, the quality of its probability estimates. Finally, we illustrate how our statistical perspective facilitates novel applications of distillation to bipartite ranking and multiclass retrieval.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- Large Language Models Are Reasoning TeachersNamgyu Ho, Laura Schmid, Se-Young YunACL 2023 · 被引用 102 次
- Better Uncertainty Calibration via Proper Scores for Classification and BeyondSebastian G. Gruber, Florian BuettnerNeurIPS 2022 · 被引用 88 次
- What Makes a "Good" Data Augmentation in Knowledge Distillation - A Statistical PerspectiveHuan Wang, Suhas Lohit, Michael J. Jones, Yun FuNeurIPS 2022 · 被引用 62 次
- What Knowledge Gets Distilled in Knowledge Distillation?Utkarsh Ojha, Yuheng Li, Anirudh Sundara Rajan, Yingyu Liang 等NeurIPS 2023 · 被引用 58 次
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li 等NeurIPS 2022 · 被引用 46 次
它引用的顶会 Paper11
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Does label smoothing mitigate label noise?Michal Lukasik, Srinadh Bhojanapalli, Aditya Krishna Menon, Sanjiv KumarICML 2020 · 被引用 411 次
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 被引用 298 次
- Rethinking Bias-Variance Trade-off for Generalization of Neural NetworksZitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt 等ICML 2020 · 被引用 219 次
相关 Paper
- SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and GuidelinesItai Morad, Nir Shlezinger, Yonina C. EldarICLR 2026 · 被引用 1 次
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi 等NeurIPS 2021 · 被引用 318 次
- PLD: A Choice-Theoretic List-Wise Knowledge DistillationEjafa Bassam, Dawei Zhu, Kaigui BianNeurIPS 2025
- Knowledge Distillation with Perturbed Loss: From a Vanilla Teacher to a Proxy TeacherRongzhi Zhang, Jiaming Shen, Tianqi Liu, Jialu Liu 等KDD 2024 · 被引用 3 次
- On student-teacher deviations in distillation: does it pay to disobey?Vaishnavh Nagarajan, Aditya Krishna Menon, Srinadh Bhojanapalli, Hossein Mobahi 等NeurIPS 2023 · 被引用 25 次
