Logit Standardization in Knowledge Distillation
Shangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang, Xiaochun Cao
摘要
Knowledge distillation involves transferring soft labels from a teacher to a student using a shared temperature-based softmax function. However, the assumption of a shared temperature between teacher and student implies a mandatory exact match between their logits in terms of logit range and variance. This side-effect limits the performance of student, considering the capacity discrepancy between them and the finding that the innate logit relations of teacher are sufficient for student to learn. To address this issue, we propose setting the temperature as the weighted standard deviation of logit and performing a plug-and-play Z-score pre-process of logit standardization before applying softmax and Kullback-Leibler divergence. Our pre-process enables student to focus on essential logit relationsfrom teacher rather than requiring a magnitude match, and can improve the performance of existing logit-based distillation methods. We also show a typical case where the conventional setting of sharing temperature between teacher and student cannot reliably yield the authentic dis-tillation evaluation; nonetheless, this challenge is success-fully alleviated by our Z-score. We extensively evaluate our method for various student and teacher models on CIFAR-100 and ImageNet, showing its significant superiority. The vanilla knowledge distillation powered by our pre-process can achieve favorable performance against state-of-the-art methods, and other distillation variants can obtain considerable gain with the assistance of our pre-process. The codes, pre-trained models and logs are released on Github.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper62
- Knowledge Distillation with Refined LogitsWujie Sun, Defang Chen, Siwei Lyu, Genlang Chen 等ICCV 2025 · 被引用 13 次
- Cross-View Consistency Regularisation for Knowledge DistillationWeijia Zhang, Dongnan Liu, Weidong Cai, Chao MaACM MM 2024 · 被引用 11 次
- Local Dense Logit Relations for Enhanced Knowledge DistillationLiuchi Xu, Kang Liu, Jinshuai Liu, Lu Wang 等ICCV 2025 · 被引用 10 次
- Maintaining Fairness in Logit-based Knowledge Distillation for Class-Incremental LearningZijian Gao, Shanhao Han, Xingxing Zhang, Kele Xu 等AAAI 2025 · 被引用 10 次
- A Good Teacher Adapts Their Knowledge for DistillationChengyao Qian, Trung Le, Mehrtash HarandiICCV 2025 · 被引用 8 次
它引用的顶会 Paper23
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
相关 Paper
- Knowledge Distillation Based on Transformed Teacher MatchingKaixiang Zheng, En-Hui YangICLR 2024 · 被引用 40 次
- Enhancing Logits Distillation with Plug&Play Kendall's τ Ranking LossYuchen Guan, Runxi Cheng, Kang Liu, Chun YuanICML 2025
- Curriculum Temperature for Knowledge DistillationZheng Li, Xiang Li, Lingfeng Yang, Borui Zhao 等AAAI 2023 · 被引用 277 次
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li 等NeurIPS 2022 · 被引用 46 次
- Rethinking the Dark Knowledge and Kullback-Leibler Divergence Loss in Knowledge Distillation Under Capacity MismatchingYingchao Wang, Wenqi Niu, Xingshan Yao, Li You 等AAAI 2026
