Toward Student-oriented Teacher Network Training for Knowledge Distillation
Chengyu Dong, Liyuan Liu, Jingbo Shang
Abstract
How to conduct teacher training for knowledge distillation is still an open problem. It has been widely observed that a best-performing teacher does not necessarily yield the best-performing student, suggesting a fundamental discrepancy between the current teacher training practice and the ideal teacher training strategy. To fill this gap, we explore the feasibility of training a teacher that is oriented toward student performance with empirical risk minimization (ERM). Our analyses are inspired by the recent findings that the effectiveness of knowledge distillation hinges on the teacher's capability to approximate the true label distribution of training inputs. We theoretically establish that the ERM minimizer can approximate the true label distribution of training data as long as the feature extractor of the learner network is Lipschitz continuous and is robust to feature transformations. In light of our theory, we propose a teacher training method SoTeacher which incorporates Lipschitz regularization and consistency regularization into ERM. Experiments on benchmark datasets using various knowledge distillation algorithms and teacher-student pairs confirm that SoTeacher can improve student accuracy consistently.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8a67384b-b835-4c48-ac69-d0e8b004389aCited by top-tier papers4
- Reliable Model Watermarking: Defending against Theft without Compromising on EvasionHongyu Zhu, Sichu Liang, Wentao Hu, Fangqi Li et al.ACM MM 2024 · 14 citations
- SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and GuidelinesItai Morad, Nir Shlezinger, Yonina C. EldarICLR 2026 · 1 citation
- Differentiable JPEG-based Input Perturbation for Knowledge Distillation Amplification via Conditional Mutual Information MaximizationSIYU CHEN, Kaixiang Zheng, Ahmed H. Salamah, EN-HUI YANGICLR 2026
- Gap Preserving Distillation by Building Bidirectional Mappings with A Dynamic TeacherYong Guo, Shulian Zhang, Haolin Pan, Jing Liu et al.ICLR 2025
Builds on17
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph et al.ICLR 2020 · 1,572 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
Related papers
- Lipschitz Continuity Guided Knowledge DistillationYuzhang Shang, Bin Duan, Ziliang Zong, Liqiang Nie et al.ICCV 2021 · 31 citations
- Knowledge Distillation with Perturbed Loss: From a Vanilla Teacher to a Proxy TeacherRongzhi Zhang, Jiaming Shen, Tianqi Liu, Jialu Liu et al.KDD 2024 · 3 citations
- A statistical perspective on distillationAditya Krishna Menon, Ankit Singh Rawat, Sashank J. Reddi, Seungyeon Kim et al.ICML 2021 · 97 citations
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi et al.NeurIPS 2021 · 318 citations
- Learning From Biased Soft LabelsHua Yuan, Yu Shi, Ning Xu, Xu Yang et al.NeurIPS 2023 · 16 citations
