Revisit the Essence of Distilling Knowledge through Calibration
Wen-Shu Fan, Su Lu, Xin-Chun Li, De-Chuan Zhan, Le Gan
Abstract
Knowledge Distillation (KD) has evolved into a practical technology for transferring knowledge from a well-performing model (teacher) to a weak model (student). A counter-intuitive phenomenon known as capacity mismatch has been identified, wherein KD performance may not be good when a better teacher instructs the student. Various preliminary methods have been proposed to alleviate capacity mismatch, but a unifying explanation for its cause still lacks. In this paper, we propose a unifying analytical framework to pinpoint the core of capacity mismatch based on calibration. Through extensive analytical experiments, we observe a positive correlation between the calibration of the teacher model and the KD performance with original KD methods. As this correlation arises due to the sensitivity of metrics (e.g., KL divergence) to calibration, we recommend employing measurements insensitive to calibration such as ranking-based loss. Our experiments demonstrate that ranking-based loss can effectively replace KL divergence, aiding large models with poor calibration to teach better.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 533d84ec-68c2-458d-9c17-6e5b60f26e95Cited by top-tier papers5
- SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and GuidelinesItai Morad, Nir Shlezinger, Yonina C. EldarICLR 2026 · 1 citation
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li et al.ACL 2025
- PLD: A Choice-Theoretic List-Wise Knowledge DistillationEjafa Bassam, Dawei Zhu, Kaigui BianNeurIPS 2025
- Rethinking the Dark Knowledge and Kullback-Leibler Divergence Loss in Knowledge Distillation Under Capacity MismatchingYingchao Wang, Wenqi Niu, Xingshan Yao, Li You et al.AAAI 2026
- Distillation Scaling LawsDan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram et al.ICML 2025
Builds on12
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 741 citations
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian et al.NeurIPS 2022 · 477 citations
Related papers
- A Good Teacher Adapts Their Knowledge for DistillationChengyao Qian, Trung Le, Mehrtash HarandiICCV 2025 · 8 citations
- Enhancing Logits Distillation with Plug&Play Kendall's τ Ranking LossYuchen Guan, Runxi Cheng, Kang Liu, Chun YuanICML 2025
- Knowledge Distillation with Perturbed Loss: From a Vanilla Teacher to a Proxy TeacherRongzhi Zhang, Jiaming Shen, Tianqi Liu, Jialu Liu et al.KDD 2024 · 3 citations
- Evaluation-oriented Knowledge Distillation for Deep Face RecognitionYuge Huang, Jiaxiang Wu, Xingkun Xu, Shouhong DingCVPR 2022 · 35 citations
- Knowledge Distillation with Auxiliary VariableBo Peng, Zhen Fang, Guangquan Zhang, Jie LuICML 2024 · 7 citations
