Knowledge Distillation as Semiparametric Inference
Tri Dao, Govinda M. Kamath, Vasilis Syrgkanis, Lester Mackey
摘要
A popular approach to model compression is to train an inexpensive student model to mimic the class probabilities of a highly accurate but cumbersome teacher model. Surprisingly, this two-step knowledge distillation process often leads to higher accuracy than training the student directly on labeled data. To explain and enhance this phenomenon, we cast knowledge distillation as a semiparametric inference problem with the optimal student model as the target, the unknown Bayes class probabilities as nuisance, and the teacher probabilities as a plug-in nuisance estimate. By adapting modern semiparametric tools, we derive new guarantees for the prediction error of standard distillation and develop two enhancements -- cross-fitting and loss correction -- to mitigate the impact of teacher overfitting and underfitting on student performance. We validate our findings empirically on both tabular and image data and observe consistent improvements from our knowledge distillation enhancements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- LightTS: Lightweight Time Series Classification with Adaptive Ensemble DistillationDavid Campos, Miao Zhang, Bin Yang, Tung Kieu 等SIGMOD 2023 · 被引用 105 次
- A statistical perspective on distillationAditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim 等ICML 2021 · 被引用 97 次
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li 等NeurIPS 2022 · 被引用 46 次
- Bayes Conditional Distribution Estimation for Knowledge Distillation Based on Conditional Mutual InformationLinfeng Ye, Shayan Mohajer Hamidi, Renhao Tan, En-Hui YangICLR 2024 · 被引用 27 次
- On student-teacher deviations in distillation: does it pay to disobey?Vaishnavh Nagarajan, Aditya Krishna Menon, Srinadh Bhojanapalli, Hossein Mobahi 等NeurIPS 2023 · 被引用 25 次
它引用的顶会 Paper8
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Adversarially Robust DistillationMicah Goldblum, Liam Fowl, Soheil Feizi, Tom GoldsteinAAAI 2020 · 被引用 258 次
- And the Bit Goes Down: Revisiting the Quantization of Neural NetworksPierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham 等ICLR 2020 · 被引用 157 次
- Fast, Accurate, and Simple Models for Tabular Data via Augmented DistillationRasool Fakoor, Jonas Mueller, Nick Erickson, Pratik Chaudhari 等NeurIPS 2020 · 被引用 65 次
相关 Paper
- Knowledge Distillation with Refined LogitsWujie Sun, Defang Chen, Siwei Lyu, Genlang Chen 等ICCV 2025 · 被引用 13 次
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 被引用 66 次
- SLaM: Student-Label Mixing for Distillation with Unlabeled ExamplesVasilis Kontonis, Fotis Iliopoulos, Khoa Trinh, Cenk Baykal 等NeurIPS 2023 · 被引用 10 次
- Comprehensive Knowledge Distillation with Causal InterventionXiang Deng, Zhongfei ZhangNeurIPS 2021 · 被引用 44 次
- PLD: A Choice-Theoretic List-Wise Knowledge DistillationEjafa Bassam, Dawei Zhu, Kaigui BianNeurIPS 2025
