Revisit the Essence of Distilling Knowledge through Calibration
Wen-Shu Fan, Su Lu, Xin-Chun Li, De-Chuan Zhan, Le Gan
摘要
Knowledge Distillation (KD) has evolved into a practical technology for transferring knowledge from a well-performing model (teacher) to a weak model (student). A counter-intuitive phenomenon known as capacity mismatch has been identified, wherein KD performance may not be good when a better teacher instructs the student. Various preliminary methods have been proposed to alleviate capacity mismatch, but a unifying explanation for its cause still lacks. In this paper, we propose a unifying analytical framework to pinpoint the core of capacity mismatch based on calibration. Through extensive analytical experiments, we observe a positive correlation between the calibration of the teacher model and the KD performance with original KD methods. As this correlation arises due to the sensitivity of metrics (e.g., KL divergence) to calibration, we recommend employing measurements insensitive to calibration such as ranking-based loss. Our experiments demonstrate that ranking-based loss can effectively replace KL divergence, aiding large models with poor calibration to teach better.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- SGD-Based Knowledge Distillation with Bayesian Teachers: Theory and GuidelinesItai Morad, Nir Shlezinger, Yonina C. EldarICLR 2026 · 被引用 1 次
- Maximizing the Effectiveness of Larger BERT Models for CompressionWen-Shu Fan, Su Lu, Shangyu Xing, Xin-Chun Li 等ACL 2025
- PLD: A Choice-Theoretic List-Wise Knowledge DistillationEjafa Bassam, Dawei Zhu, Kaigui BianNeurIPS 2025
- Rethinking the Dark Knowledge and Kullback-Leibler Divergence Loss in Knowledge Distillation Under Capacity MismatchingYingchao Wang, Wenqi Niu, Xingshan Yao, Li You 等AAAI 2026
- Distillation Scaling LawsDan Busbridge, Amitis Shidani, Floris Weers, Jason Ramapuram 等ICML 2025
它引用的顶会 Paper12
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian 等NeurIPS 2022 · 被引用 477 次
相关 Paper
- A Good Teacher Adapts Their Knowledge for DistillationChengyao Qian, Trung Le, Mehrtash HarandiICCV 2025 · 被引用 8 次
- Enhancing Logits Distillation with Plug&Play Kendall's τ Ranking LossYuchen Guan, Runxi Cheng, Kang Liu, Chun YuanICML 2025
- Knowledge Distillation with Perturbed Loss: From a Vanilla Teacher to a Proxy TeacherRongzhi Zhang, Jiaming Shen, Tianqi Liu, Jialu Liu 等KDD 2024 · 被引用 3 次
- Evaluation-oriented Knowledge Distillation for Deep Face RecognitionYuge Huang, Jiaxiang Wu, Xingkun Xu, Shouhong DingCVPR 2022 · 被引用 35 次
- Knowledge Distillation with Auxiliary VariableBo Peng, Zhen Fang, Guangquan Zhang, Jie LuICML 2024 · 被引用 7 次
