Multi-Level Logit Distillation
Ying Jin, Jiaqi Wang, Dahua Lin
摘要
Knowledge Distillation (KD) aims at distilling the knowledge from the large teacher model to a lightweight student model. Mainstream KD methods can be divided into two categories, logit distillation, and feature distillation. The former is easy to implement, but inferior in performance, while the latter is not applicable to some practical circumstances due to concerns such as privacy and safety. Towards this dilemma, in this paper, we explore a stronger logit distillation method via making better utilization of logit outputs. Concretely, we propose a simple yet effective approach to logit distillation via multilevel prediction alignment. Through this framework, the prediction alignment is not only conducted at the instance level, but also at the batch and class level, through which the student model learns instance prediction, input correlation, and category correlation simultaneously. In addition, a prediction augmentation mechanism based on model calibration further boosts the performance. Extensive experiment results validate that our method enjoys consistently higher performance than previous logit distillation methods, and even reaches competitive performance with mainstream feature distillation methods. Code is available at https://github.com/Jin-Ying/Multi-Level-Logit-Distillation .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper44
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang 等CVPR 2024 · 被引用 183 次
- Wasserstein Distance Rivals Kullback-Leibler Divergence for Knowledge DistillationJiaming Lv, Haoyuan Yang, Peihua LiNeurIPS 2024 · 被引用 59 次
- DistilVPR: Cross-Modal Knowledge Distillation for Visual Place RecognitionSijie Wang, Rui She, Qiyu Kang, Xingchao Jian 等AAAI 2024 · 被引用 14 次
- Knowledge Distillation with Refined LogitsWujie Sun, Defang Chen, Siwei Lyu, Genlang Chen 等ICCV 2025 · 被引用 13 次
- Cross-View Consistency Regularisation for Knowledge DistillationWeijia Zhang, Dongnan Liu, Weidong Cai, Chao MaACM MM 2024 · 被引用 11 次
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine 等AAAI 2020 · 被引用 1,361 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 被引用 1,214 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
相关 Paper
- Scale Decoupled DistillationShicai Wei, Chunbo Luo, Yang LuoCVPR 2024 · 被引用 32 次
- Streamlined Knowledge DistillationHyeon-Jin Jeong, Han-Jin Lee, Seok-Hwan ChoiCVPR 2026 · 被引用 1 次
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian 等NeurIPS 2022 · 被引用 477 次
- Grouped Knowledge Distillation for Deep Face RecognitionWeisong Zhao, Xiangyu Zhu, Kaiwen Guo, Xiaoyu Zhang 等AAAI 2023 · 被引用 12 次
- Multi-Label Knowledge DistillationPenghui Yang, Ming-Kun Xie, Chen-Chen Zong, Lei Feng 等ICCV 2023 · 被引用 16 次
