Knowledge Distillation Performs Partial Variance Reduction
Mher Safaryan, Alexandra Peste, Dan Alistarh
Abstract
Knowledge distillation is a popular approach for enhancing the performance of ''student'' models, with lower representational capacity, by taking advantage of more powerful ''teacher'' models. Despite its apparent simplicity and widespread use, the underlying mechanics behind knowledge distillation (KD) are still not fully understood. In this work, we shed new light on the inner workings of this method, by examining it from an optimization perspective. We show that, in the context of linear and deep linear models, KD can be interpreted as a novel type of stochastic variance reduction mechanism. We provide a detailed convergence analysis of the resulting dynamics, which hold under standard assumptions for both strongly-convex and non-convex losses, showing that KD acts as a form of partial variance reduction, which can reduce the stochastic gradient noise, but may not eliminate it completely, depending on the properties of the ''teacher'' model. Our analysis puts further emphasis on the need for careful parametrization of KD, in particular w.r.t. the weighting of the distillation loss, and is validated empirically on both linear models and deep neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- InfiGFusion: Graph-on-Logits Distillation via Efficient Gromov-Wasserstein for Model FusionYuanyi Wang, Zhaoyi Yan, Yiming Zhang, Qi Zhou et al.NeurIPS 2025 · 12 citations
- Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time ScalingSachin Goyal, David Lopez-Paz, Kartik AhujaICLR 2026 · 11 citations
- Debiased Distillation for Consistency RegularizationLu Wang, Liuchi Xu, Xiong Yang, Zhenhua Huang et al.AAAI 2025 · 6 citations
- Prediction-Powered Semi-Supervised Learning with Online Power TuningNoa Shoham, Ron Dorfman, Shalev Shaer, Kfir Y. Levy et al.NeurIPS 2025 · 5 citations
- SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and GuidelinesItai Morad, Nir Shlezinger, Yonina C. EldarICLR 2026 · 1 citation
Builds on8
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited DevicesZhiqing Sun, Hongkun Yu, Xiaodan Song, Renjie Liu et al.ACL 2020 · 660 citations
- Preserved central model for faster bidirectional compression in distributed settingsConstantin Philippenko, Aymeric DieuleveutNeurIPS 2021 · 37 citations
- Smoothness Matrices Beat Smoothness Constants: Better Communication Compression Techniques for Distributed OptimizationMher Safaryan, Filip Hanzely, Peter RichtárikNeurIPS 2021 · 32 citations
Related papers
- What Makes a Strong Model? A Unified Spectral Analysis of Knowledge Transfer over High-dimensional Linear RegressionWendao Wu, Fangqing Zhang, Haihan Zhang, Cong FangICML 2026
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang et al.CVPR 2020
- Knowledge Diffusion for DistillationTao Huang, Yuan Zhang, Mingkai Zheng, Shan You et al.NeurIPS 2023 · 125 citations
- Learning Student-Friendly Teacher Networks for Knowledge DistillationDae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim et al.NeurIPS 2021 · 134 citations
- Bayesian Knowledge Distillation: A Bayesian Perspective of Distillation with Uncertainty QuantificationLuyang Fang, Yongkai Chen, Wenxuan Zhong, Ping MaICML 2024 · 10 citations
