Knowledge Distillation with the Reused Teacher Classifier
Defang Chen, Jian-Ping Mei, Hailin Zhang, Can Wang, Yan Feng, Chun Chen
摘要
Knowledge distillation aims to compress a powerful yet cumbersome teacher model into a lightweight student model without much sacrifice of performance. For this purpose, various approaches have been proposed over the past few years, generally with elaborately designed knowledge rep-resentations, which in turn increase the difficulty of model development and interpretation. In contrast, we empirically show that a simple knowledge distillation technique is enough to significantly narrow down the teacher-student performance gap. We directly reuse the discriminative classifier from the pre-trained teacher model for student inference and train a student encoder through feature alignment with a single ℓ <inf xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</inf> loss. In this way, the student model is able to achieve exactly the same performance as the teacher model provided that their extracted features are perfectly aligned. An additional projector is developed to help the student encoder match with the teacher classifier, which renders our technique applicable to various teacher and student architectures. Extensive experiments demonstrate that our technique achieves state-of-the-art results at the modest cost of compression ratio due to the added projector.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper52
- Logit Standardization in Knowledge DistillationShangquan Sun, Wenqi Ren, Jingzhi Li, Rui Wang 等CVPR 2024 · 被引用 183 次
- Encoding Time-Series Explanations through Self-Supervised Model Behavior ConsistencyOwen Queen, Tom Hartvigsen, Teddy Koker, Huan He 等NeurIPS 2023 · 被引用 55 次
- Scale Decoupled DistillationShicai Wei, Chunbo Luo, Yang LuoCVPR 2024 · 被引用 32 次
- Distribution Shift Matters for Knowledge Distillation with Webly Collected ImagesJialiang Tang, Shuo Chen, Gang Niu, Masashi Sugiyama 等ICCV 2023 · 被引用 21 次
- ScaleKD: Strong Vision Transformers Could Be Excellent TeachersJiawei Fan, Chao Li, Xiaolong Liu, Anbang YaoNeurIPS 2024 · 被引用 20 次
它引用的顶会 Paper18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Do We Really Need to Access the Source Data? Source Hypothesis Transfer for Unsupervised Domain AdaptationJian Liang, Dapeng Hu, Jiashi FengICML 2020 · 被引用 1,624 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
相关 Paper
- Knowledge distillation via softmax regression representation learningJing Yang, Brais Martínez, Adrian Bulat, Georgios TzimiropoulosICLR 2021 · 被引用 55 次
- Improved Feature Distillation via Projector EnsembleYudong Chen, Sen Wang, Jiajun Liu, Xuwei Xu 等NeurIPS 2022 · 被引用 73 次
- Knowledge Diffusion for DistillationTao Huang, Yuan Zhang, Mingkai Zheng, Shan You 等NeurIPS 2023 · 被引用 125 次
- KD-Zero: Evolving Knowledge Distiller for Any Teacher-Student PairsLujun Li, Peijie Dong, Anggeng Li, Zimian Wei 等NeurIPS 2023 · 被引用 49 次
- UniKD: Universal Knowledge Distillation for Mimicking Homogeneous or Heterogeneous Object DetectorsShanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu 等ICCV 2023 · 被引用 7 次
