Knowledge Distillation Performs Partial Variance Reduction
Mher Safaryan, Alexandra Peste, Dan Alistarh
摘要
Knowledge distillation is a popular approach for enhancing the performance of ''student'' models, with lower representational capacity, by taking advantage of more powerful ''teacher'' models. Despite its apparent simplicity and widespread use, the underlying mechanics behind knowledge distillation (KD) are still not fully understood. In this work, we shed new light on the inner workings of this method, by examining it from an optimization perspective. We show that, in the context of linear and deep linear models, KD can be interpreted as a novel type of stochastic variance reduction mechanism. We provide a detailed convergence analysis of the resulting dynamics, which hold under standard assumptions for both strongly-convex and non-convex losses, showing that KD acts as a form of partial variance reduction, which can reduce the stochastic gradient noise, but may not eliminate it completely, depending on the properties of the ''teacher'' model. Our analysis puts further emphasis on the need for careful parametrization of KD, in particular w.r.t. the weighting of the distillation loss, and is validated empirically on both linear models and deep neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model FusionYuanyi Wang, Zhaoyi Yan, Yiming Zhang, Qi Zhou 等NeurIPS 2025 · 被引用 12 次
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time ScalingSachin Goyal, David Lopez-Paz, Kartik AhujaICLR 2026 · 被引用 11 次
- Debiased Distillation for Consistency RegularizationLu Wang, Liuchi Xu, Xiong Yang, Zhenhua Huang 等AAAI 2025 · 被引用 6 次
- Prediction-Powered Semi-Supervised Learning with Online Power TuningNoa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy 等NeurIPS 2025 · 被引用 5 次
- SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and GuidelinesItai Morad, Nir Shlezinger, Yonina C. EldarICLR 2026 · 被引用 1 次
它引用的顶会 Paper8
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu 等ACL 2020 · 被引用 660 次
- Preserved central model for faster bidirectional compression in distributed settingsConstantin Philippenko, Aymeric DieuleveutNeurIPS 2021 · 被引用 37 次
- Smoothness Matrices Beat Smoothness Constants: Better Communication Compression Techniques for Distributed OptimizationMher Safaryan, Filip Hanzely, Peter RichtárikNeurIPS 2021 · 被引用 32 次
相关 Paper
- What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear RegressionWendao Wu, Fangqing Zhang, Haihan Zhang, Cong FangICML 2026
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang 等CVPR 2020
- Knowledge Diffusion for DistillationTao Huang, Yuan Zhang, Mingkai Zheng, Shan You 等NeurIPS 2023 · 被引用 125 次
- Learning Student-Friendly Teacher Networks for Knowledge DistillationDae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim 等NeurIPS 2021 · 被引用 134 次
- Bayesian Knowledge Distillation: A Bayesian Perspective of Distillation with Uncertainty QuantificationLuyang Fang, Yongkai Chen, Wenxuan Zhong, Ping MaICML 2024 · 被引用 10 次
