SDPGO: Efficient Self-Distillation Training Meets Proximal Gradient Optimization
Tongtong Su, Yun Liao, Fengbo Zheng
Abstract
Self-knowledge distillation (SKD) enables single-model training by distilling knowledge from the model’s own output, eliminating the need for a separate teacher network required in conventional distillation methods. However, current SKD meth-ods focus mainly on replicating common features in the student model, neglecting the extraction of key features that significantly enhance student learning. Inspired by this, we devise a self-knowledge distillation framework entitled Self-Distillation training via Proximal Gradient Optimization or SDPGO , which utilizes gradient information to identify and assign greater weight to features that significantly impact classification performance, enabling the network to learn the most relevant features during training. Specifically, the proposed framework refines the gradient information into a dynamically changing weighting factor to evaluate the distillation knowledge via the dynamic weight adjustment scheme. Meanwhile, we devise the sequential iterative learning module to dynamically optimize knowledge transfer by leveraging historical predictions and real-time gradients, stabilizing training through mini-batch-based KL divergence refinement while adaptively prioritizing task-critical features for efficient self-distillation. Comprehensive experiments on image classification, object detection, and semantic segmentation demonstrate that our method consistently surpasses recent state-of-the-art knowledge distillation techniques. Code is available at: https://github.com/nanxiaotong/SDGPO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4405f4fd-7128-40ce-bd95-082bf9e4c5f7Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
Related papers
- Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge DistillationMingi Ji, Seungjae Shin, Seunghyun Hwang, Gibeom Park et al.CVPR 2021
- Self-Decoupling and Ensemble Distillation for Efficient SegmentationYuang Liu, Wei Zhang, Jun WangAAAI 2023 · 4 citations
- Multi-Teacher Knowledge Distillation with Reinforcement Learning for Visual RecognitionChuanguang Yang, Xinqiang Yu, Han Yang, Zhulin An et al.AAAI 2025 · 26 citations
- Student-Oriented Teacher Knowledge Refinement for Knowledge DistillationChaomin Shen, Yaomin Huang, Haokun Zhu, Jinsong Fan et al.ACM MM 2024 · 2 citations
- Localization Distillation for Dense Object DetectionZhaohui Zheng, Rongguang Ye, Ping Wang, Dongwei Ren et al.CVPR 2022 · 177 citations
