Self-Distillation from the Last Mini-Batch for Consistency Regularization
Yiqing Shen, Liwu Xu, Yuzhe Yang, Yaqian Li, Yandong Guo
摘要
Knowledge distillation (KD) shows a bright promise as a powerful regularization strategy to boost generalization ability by leveraging learned sample-level soft targets. Yet, employing a complex pre-trained teacher network or an ensemble of peer students in existing KD is both timeconsuming and computationally costly. Various self KD methods have been proposed to achieve higher distillation efficiency. However, they either require extra network architecture modification or are difficult to parallelize. To cope with these challenges, we propose an efficient and reliable self-distillation framework, named Self-Distillation from Last Mini-Batch (DLB). Specifically, we rearrange the sequential sampling by constraining half of each mini-batch coinciding with the previous iteration. Meanwhile, the rest half will coincide with the upcoming iteration. Afterwards, the former half mini-batch distills on-the-fly soft targets generated in the previous iteration. Our proposed mechanism guides the training stability and consistency, resulting in robustness to label noise. Moreover, our method is easy to implement, without taking up extra run-time memory or requiring model structure modification. Experimental results on three classification benchmarks illustrate that our approach can consistently outperform state-of-the-art self-distillation approaches with different network architectures. Additionally, our method shows strong compatibility with augmentation strategies by gaining additional performance improvement. The code is available at https: //github.com/Meta-knowledge-Lab/DLB .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Distilling Reliable Knowledge for Instance-Dependent Partial Label LearningDong-Dong Wu, Deng-Bao Wang, Min-Ling ZhangAAAI 2024 · 被引用 16 次
- Bayesian Knowledge Distillation: A Bayesian Perspective of Distillation with Uncertainty QuantificationLuyang Fang, Yongkai Chen, Wenxuan Zhong, Ping MaICML 2024 · 被引用 10 次
- EPSD: Early Pruning with Self-Distillation for Efficient Model CompressionDong Chen, Ning Liu, Yichen Zhu, Zhengping Che 等AAAI 2024 · 被引用 9 次
- Boosting Residual Networks with Group KnowledgeShengji Tang, Peng Ye, Baopu Li, Weihao Lin 等AAAI 2024 · 被引用 7 次
- Towards Robust and Generalized Parameter-Efficient Fine-Tuning for Noisy Label LearningYeachan Kim, Junho Kim, SangKeun LeeACL 2024 · 被引用 4 次
它引用的顶会 Paper12
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh 等ICCV 2019 · 被引用 5,843 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
- Online Knowledge Distillation with Diverse PeersDefang Chen, Jian-Ping Mei, Can Wang, Yan Feng 等AAAI 2020 · 被引用 354 次
- Self-Knowledge Distillation with Progressive Refinement of TargetsKyungyul Kim, Byeongmoon Ji, Doyoung Yoon, Sangheum HwangICCV 2021 · 被引用 251 次
- Self-supervised Label Augmentation via Input TransformationsHankook Lee, Sung Ju Hwang, Jinwoo ShinICML 2020 · 被引用 218 次
相关 Paper
- Self-Decoupling and Ensemble Distillation for Efficient SegmentationYuang Liu, Wei Zhang, Jun WangAAAI 2023 · 被引用 4 次
- Revisiting Knowledge Distillation via Label Smoothing RegularizationLi Yuan, Francis E. H. Tay, Guilin Li, Tao Wang 等CVPR 2020
- DA-KD: Difficulty-Aware Knowledge Distillation for Efficient Large Language ModelsChangyi He, Yifu Ding, Jinyang Guo, Ruihao Gong 等ICML 2025
- Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge DistillationMingi Ji, Seungjae Shin, Seunghyun Hwang, Gibeom Park 等CVPR 2021
- Adaptive Dual Guidance Knowledge DistillationTong Li, Long Liu, Kang Liu, Xin Wang 等AAAI 2025 · 被引用 1 次
