Self-Knowledge Distillation with Progressive Refinement of Targets
Kyungyul Kim, Byeongmoon Ji, Doyoung Yoon, Sangheum Hwang
摘要
The generalization capability of deep neural networks has been substantially improved by applying a wide spectrum of regularization methods, e.g., restricting function space, injecting randomness during training, augmenting data, etc. In this work, we propose a simple yet effective regularization method named progressive self-knowledge distillation (PS-KD), which progressively distills a model's own knowledge to soften hard targets (i.e., one-hot vectors) during training. Hence, it can be interpreted within a framework of knowledge distillation as a student becomes a teacher itself. Specifically, targets are adjusted adaptively by combining the ground-truth and past predictions from the model itself. We show that PS-KD provides an effect of hard example mining by rescaling gradients according to difficulty in classifying examples. The proposed method is applicable to any supervised learning tasks with hard targets and can be easily combined with existing regularization methods to further enhance the generalization performance. Furthermore, it is confirmed that PS-KD achieves not only better accuracy, but also provides high quality of confidence estimates in terms of calibration as well as ordinal ranking. Extensive experimental results on three different tasks, image classification, object detection, and machine translation, demonstrate that our method consistently improves the performance of the state-of-the-art baselines. The code is available at https://github.com/lgcnsai/PS-KD-Pytorch .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper45
- Single-Domain Generalized Object Detection in Urban Scene via Cyclic-Disentangled Self-DistillationAming Wu, Cheng DengCVPR 2022 · 被引用 110 次
- Self-Distillation from the Last Mini-Batch for Consistency RegularizationYiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li 等CVPR 2022 · 被引用 88 次
- DINE: Domain Adaptation from Single and Multiple Black-box PredictorsJian Liang, Dapeng Hu, Jiashi Feng, Ran HeCVPR 2022 · 被引用 80 次
- Introspective Distillation for Robust Question AnsweringYulei Niu, Hanwang ZhangNeurIPS 2021 · 被引用 74 次
- Cross-Task Knowledge Distillation in Multi-Task RecommendationChenxiao Yang, Junwei Pan, Xiaofeng Gao, Tingyu Jiang 等AAAI 2022 · 被引用 58 次
它引用的顶会 Paper5
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Confidence-Aware Learning for Deep Neural NetworksJooyoung Moon, Jihyo Kim, Younghak Shin, Sangheum HwangICML 2020 · 被引用 184 次
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang 等CVPR 2020
- Regularizing Class-Wise Predictions via Self-Knowledge DistillationSukmin Yun, Jongjin Park, Kimin Lee, Jinwoo ShinCVPR 2020
相关 Paper
- Strengthen Out-of-Distribution Detection Capability with Progressive Self-Knowledge DistillationYang Yang, Haonan XuICML 2025
- FerKD: Surgical Label Adaptation for Efficient DistillationZhiqiang ShenICCV 2023 · 被引用 7 次
- Adaptive Label Smoothing with Self-Knowledge in Natural Language GenerationDongkyu Lee, Ka Chun Cheung, Nevin L. ZhangEMNLP 2022 · 被引用 5 次
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu 等CVPR 2022 · 被引用 835 次
- From Knowledge Distillation to Self-Knowledge Distillation: A Unified Approach with Normalized Loss and Customized Soft LabelsZhendong Yang, Ailing Zeng, Zhe Li, Tianke Zhang 等ICCV 2023 · 被引用 141 次
