Revisiting Knowledge Distillation: An Inheritance and Exploration Framework
Zhen Huang, Xu Shen, Jun Xing, Tongliang Liu, Xinmei Tian, Houqiang Li, Bing Deng, Jianqiang Huang, Xian-Sheng Hua
摘要
Knowledge Distillation (KD) is a popular technique to transfer knowledge from a teacher model or ensemble to a student model. Its success is generally attributed to the privileged information on similarities/consistency between the class distributions or intermediate feature representations of the teacher model and the student model. However, directly pushing the student model to mimic the probabilities/features of the teacher model to a large extent limits the student model in learning undiscovered knowledge/features. In this paper, we propose a novel inheritance and exploration knowledge distillation framework (IE-KD), in which a student model is split into two parts -inheritance and exploration. The inheritance part is learned with a similarity loss to transfer the existing learned knowledge from the teacher model to the student model, while the exploration part is encouraged to learn representations different from the inherited ones with a dis-similarity loss. Our IE-KD framework is generic and can be easily combined with existing distillation or mutual learning methods for training deep neural networks. Extensive experiments demonstrate that these two parts can jointly push the student model to learn more diversified and effective representations, and our IE-KD can be a general technique to improve the student network to achieve SOTA performance. Furthermore, by applying our IE-KD to the training of two networks, the performance of both can be improved w.r.t. deep mutual learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- X -Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense CaptioningZhihao Yuan, Xu Yan, Yinghong Liao, Yao Guo 等CVPR 2022 · 被引用 72 次
- Let Images Give You More: Point Cloud Cross-Modal Training for Shape AnalysisXu Yan, Heshen Zhan, Chaoda Zheng, Jiantao Gao 等NeurIPS 2022 · 被引用 49 次
- Asymmetric Temperature Scaling Makes Larger Networks Teach Well AgainXin-Chun Li, Wen-Shu Fan, Shaoming Song, Yinchuan Li 等NeurIPS 2022 · 被引用 46 次
- Roadblocks for Temporarily Disabling Shortcuts and Learning New KnowledgeHongjing Niu, Hanting Li, Feng Zhao, Bin LiNeurIPS 2022 · 被引用 9 次
- Why Knowledge Distillation Works in Generative Models: A Minimal Working ExplanationSungmin Cha, Kyunghyun ChoNeurIPS 2025 · 被引用 9 次
相关 Paper
- Complementary Relation Contrastive DistillationJinguo Zhu, Shixiang Tang, Dapeng Chen, Shijie Yu 等CVPR 2021
- Learning Student-Friendly Teacher Networks for Knowledge DistillationDae Young Park, Moon-Hyun Cha, Changwook Jeong, Daesin Kim 等NeurIPS 2021 · 被引用 134 次
- Heterogeneous Knowledge Distillation Using Information Flow ModelingNikolaos Passalis, Maria Tzelepi, Anastasios TefasCVPR 2020
- Neural Collapse Inspired Knowledge DistillationShuoxi Zhang, Zijian Song, Kun HeAAAI 2025 · 被引用 2 次
- Exploring Inter-Channel Correlation for Diversity-preserved Knowledge DistillationLi Liu, Qingle Huang, Sihao Lin, Hongwei Xie 等ICCV 2021 · 被引用 132 次
