Self-Knowledge Distillation with Progressive Refinement of Targets
Kyungyul Kim, Byeongmoon Ji, Doyoung Yoon, Sangheum Hwang
Abstract
The generalization capability of deep neural networks has been substantially improved by applying a wide spectrum of regularization methods, e.g., restricting function space, injecting randomness during training, augmenting data, etc. In this work, we propose a simple yet effective regularization method named progressive self-knowledge distillation (PS-KD), which progressively distills a model's own knowledge to soften hard targets (i.e., one-hot vectors) during training. Hence, it can be interpreted within a framework of knowledge distillation as a student becomes a teacher itself. Specifically, targets are adjusted adaptively by combining the ground-truth and past predictions from the model itself. We show that PS-KD provides an effect of hard example mining by rescaling gradients according to difficulty in classifying examples. The proposed method is applicable to any supervised learning tasks with hard targets and can be easily combined with existing regularization methods to further enhance the generalization performance. Furthermore, it is confirmed that PS-KD achieves not only better accuracy, but also provides high quality of confidence estimates in terms of calibration as well as ordinal ranking. Extensive experimental results on three different tasks, image classification, object detection, and machine translation, demonstrate that our method consistently improves the performance of the state-of-the-art baselines. The code is available at https://github.com/lgcnsai/PS-KD-Pytorch .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6290c67-5647-42da-9675-e724834bd662Cited by top-tier papers45
- Single-Domain Generalized Object Detection in Urban Scene via Cyclic-Disentangled Self-DistillationAming Wu, Cheng DengCVPR 2022 · 110 citations
- Self-Distillation from the Last Mini-Batch for Consistency RegularizationYiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li et al.CVPR 2022 · 88 citations
- DINE: Domain Adaptation from Single and Multiple Black-box PredictorsJian Liang, Dapeng Hu, Jiashi Feng, Ran HeCVPR 2022 · 80 citations
- Introspective Distillation for Robust Question AnsweringYulei Niu, Hanwang ZhangNeurIPS 2021 · 74 citations
- Cross-Task Knowledge Distillation in Multi-Task RecommendationChenxiao Yang, Junwei Pan, Xiaofeng Gao, Tingyu Jiang et al.AAAI 2022 · 58 citations
Builds on5
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- Confidence-Aware Learning for Deep Neural NetworksJooyoung Moon, Jihyo Kim, Younghak Shin, Sangheum HwangICML 2020 · 184 citations
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang et al.CVPR 2020
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
Related papers
- Strengthen Out-of-Distribution Detection Capability with Progressive Self-Knowledge DistillationYang Yang, Haonan XuICML 2025
- FerKD: Surgical Label Adaptation for Efficient DistillationZhiqiang ShenICCV 2023 · 7 citations
- Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationDongkyu Lee, Ka Chun Cheung, Nevin L. ZhangEMNLP 2022 · 5 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft LabelsZhendong Yang, Ailing Zeng, Zhe Li, Tianke Zhang et al.ICCV 2023 · 141 citations
