Teach Less, Learn More: On the Undistillable Classes in Knowledge Distillation
Yichen Zhu, Ning Liu, Zhiyuan Xu, Xin Liu, Weibin Meng, Louis Wang, Zhicai Ou, Jian Tang
Abstract
Knowledge distillation (KD) can effectively compress neural networks by training a smaller network (student) to simulate the behavior of a larger one (teacher). A counter-intuitive observation is that a more expansive teacher does not make a better student, but the reasons for this phenomenon remain unclear. In this paper, we demonstrate that this is directly attributed to the presence of undistillable classes : when trained with distillation, the teacher’s knowledge of some classes is incomprehensible to the student model. We observe that while KD improves the overall accuracy, it is at the cost of the model becoming inaccurate in these undistillable classes. After establishing their widespread existence in state-of-the-art distillation methods, we illustrate their correlation with the capacity gap between teacher and student models. Finally, we present a simple “Teach Less Learn More” (TLLM) framework to identify and discard the undistillable classes during training. We validate the effectiveness of our approach on multiple datasets with varying network architectures. In all settings, our proposed method is able to exceed the performance of competitive state-of-the-art techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84c7dc8b-ff96-4bb7-bc1e-c27461919aa5Cited by top-tier papers10
- Bridging Cross-task Protocol Inconsistency for Distillation in Dense Object DetectionLongrong Yang, Xianpan Zhou, Xuewei Li, Liang Qiao et al.ICCV 2023 · 51 citations
- Revisiting Ensembling in One-Shot Federated LearningYoussef Allouah, Akash Balasaheb Dhasade, Rachid Guerraoui, Nirupam Gupta et al.NeurIPS 2024 · 21 citations
- DeepKD: A Deeply Decoupled and Denoised Knowledge Distillation TrainerHaiduo Huang, Jiangcheng Song, Yadong Zhang, Pengju RenNeurIPS 2025 · 3 citations
- Any2Policy: Learning Visuomotor Policy with Any-ModalityYichen Zhu, Zhicai Ou, Feifei Feng, Jian TangNeurIPS 2024 · 3 citations
- CIFD: Controlled Information Flow to Enhance Knowledge DistillationYashas Malur Saidutta, Rakshith Sharma Srinivasa, Jaejin Cho, Ching Hua Lee et al.NeurIPS 2024 · 2 citations
Builds on26
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Improved Knowledge Distillation via Teacher AssistantSeyed-Iman Mirzadeh, Mehrdad Farajtabar, Ang Li, Nir Levine et al.AAAI 2020 · 1,361 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
Related papers
- A Good Teacher Adapts Their Knowledge for DistillationChengyao Qian, Trung Le, Mehrtash HarandiICCV 2025 · 8 citations
- DA-KD: Difficulty-Aware Knowledge Distillation for Efficient Large Language ModelsChangyi He, Yifu Ding, Jinyang Guo, Ruihao Gong et al.ICML 2025
- Analyzing the Confidentiality of Undistillable Teachers in Knowledge DistillationSouvik Kundu, Qirui Sun, Yao Fu, Massoud Pedram et al.NeurIPS 2021 · 35 citations
- Can Students Beyond the Teacher? Distilling Knowledge from Teacher's BiasJianhua Zhang, Yi Gao, Ruyu Liu, Xu Cheng et al.AAAI 2025 · 2 citations
- Student Customized Knowledge Distillation: Bridging the Gap Between Student and TeacherYichen Zhu, Yi WangICCV 2021 · 95 citations
