Regularizing Class-Wise Predictions via Self-Knowledge Distillation
Sukmin Yun, Jongjin Park, Kimin Lee, Jinwoo Shin
Abstract
Deep neural networks with millions of parameters may suffer from poor generalization due to overfitting. To mitigate the issue, we propose a new regularization method that penalizes the predictive distribution between similar samples. In particular, we distill the predictive distribution between different samples of the same label during training. This results in regularizing the dark knowledge (i.e., the knowledge on wrong predictions) of a single network (i.e., a self-knowledge distillation) by forcing it to produce more meaningful and consistent predictions in a class-wise manner. Consequently, it mitigates overconfident predictions and reduces intra-class variations. Our experimental results on various image classification tasks demonstrate that the simple yet powerful method can significantly improve not only the generalization ability but also the calibration performance of modern convolutional neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7afab54-da09-447d-a2e6-7cc79aec891dCited by top-tier papers66
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted InstancesJihoon Tack, Sangwoo Mo, Jongheon Jeong, Jinwoo ShinNeurIPS 2020 · 755 citations
- Self-Knowledge Distillation with Progressive Refinement of TargetsKyungyul Kim, Byeongmoon Ji, Doyoung Yoon, Sangheum HwangICCV 2021 · 251 citations
- From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft LabelsZhendong Yang, Ailing Zeng, Zhe Li, Tianke Zhang et al.ICCV 2023 · 141 citations
- Single-Domain Generalized Object Detection in Urban Scene via Cyclic-Disentangled Self-DistillationAming Wu, Cheng DengCVPR 2022 · 110 citations
- Semantic Concentration for Domain AdaptationShuang Li, Mixue Xie, Fangrui Lv, Chi Harold Liu et al.ICCV 2021 · 109 citations
Builds on2
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
Related papers
- Debiased Distillation for Consistency RegularizationLu Wang, Liuchi Xu, Xiong Yang, Zhenhua Huang et al.AAAI 2025 · 6 citations
- Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationDongkyu Lee, Ka Chun Cheung, Nevin L. ZhangEMNLP 2022 · 5 citations
- Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff PerspectiveHelong Zhou, Liangchen Song, Jiajie Chen, Ye Zhou et al.ICLR 2021 · 209 citations
- Knowledge Refinery: Learning from Decoupled LabelQianggang Ding, Sifan Wu, Tao Dai, Hao Sun et al.AAAI 2021 · 15 citations
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 50 citations
