Refine Myself by Teaching Myself: Feature Refinement via Self-Knowledge Distillation
Mingi Ji, Seungjae Shin, Seunghyun Hwang, Gibeom Park, Il-Chul Moon
Abstract
Knowledge distillation is a method of transferring the knowledge from a pretrained complex teacher model to a student model, so a smaller network can replace a large teacher network at the deployment stage. To reduce the necessity of training a large teacher model, the recent literatures introduced a self-knowledge distillation, which trains a student network progressively to distill its own knowledge without a pretrained teacher network. While Selfknowledge distillation is largely divided into a data augmentation based approach and an auxiliary network based approach, the data augmentation approach looses its local information in the augmentation process, which hinders its applicability to diverse vision tasks, such as semantic segmentation. Moreover, these knowledge distillation approaches do not receive the refined feature maps, which are prevalent in the object detection and semantic segmentation community. This paper proposes a novel self-knowledge distillation method, Feature Refinement via Self-Knowledge Distillation (FRSKD), which utilizes an auxiliary self-teacher network to transfer a refined knowledge for the classifier network. Our proposed method, FRSKD, can utilize both soft label and feature-map distillations for the self-knowledge distillation. Therefore, FRSKD can be applied to classification, and semantic segmentation, which emphasize preserving the local information. We demonstrate the effectiveness of FRSKD by enumerating its performance improvements in diverse tasks and benchmark datasets. The implemented code is available at https://github.com/MingiJi/FRSKD .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ea1a990-0c67-49fb-b971-176389969e6aCited by top-tier papers8
- Curriculum Temperature for Knowledge DistillationZheng Li, Xiang Li, Lingfeng Yang, Borui Zhao et al.AAAI 2023 · 277 citations
- M3AE: Multimodal Representation Learning for Brain Tumor Segmentation with Missing ModalitiesHong Liu, Dong Wei, Donghuan Lu, Jinghan Sun et al.AAAI 2023 · 101 citations
- M2SD: Multiple Mixing Self-Distillation for Few-Shot Class-Incremental LearningJinhao Lin, Ziheng Wu, Weifeng Lin, Jun Huang et al.AAAI 2024 · 13 citations
- Tail-STEAK: Improve Friend Recommendation for Tail Users via Self-Training Enhanced Knowledge DistillationYijun Ma, Chaozhuo Li, Xiao ZhouAAAI 2024 · 10 citations
- Weighted Mutual Learning with Diversity-Driven Model CompressionMiao Zhang, Li Wang, David Campos, Wei Huang et al.NeurIPS 2022 · 10 citations
Builds on9
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen et al.ICCV 2019 · 1,069 citations
- A Comprehensive Overhaul of Feature DistillationByeongho Heo, Jeesoo Kim, Sangdoo Yun, Hyojin Park et al.ICCV 2019 · 727 citations
- Learning Lightweight Lane Detection CNNs by Self Attention DistillationYuenan Hou, Zheng Ma, Chunxiao Liu, Chen Change LoyICCV 2019 · 666 citations
- Correlation Congruence for Knowledge DistillationBaoyun Peng, Xiao Jin, Dongsheng Li, Shunfeng Zhou et al.ICCV 2019 · 625 citations
Related papers
- Self-Decoupling and Ensemble Distillation for Efficient SegmentationYuang Liu, Wei Zhang, Jun WangAAAI 2023 · 4 citations
- SDPGO: Efficient Self-Distillation Training Meets Proximal Gradient OptimizationTongtong Su, Yun Liao, Fengbo ZhengNeurIPS 2025 · 1 citation
- Student-Oriented Teacher Knowledge Refinement for Knowledge DistillationChaomin Shen, Yaomin Huang, Haokun Zhu, Jinsong Fan et al.ACM MM 2024 · 2 citations
- Student Customized Knowledge Distillation: Bridging the Gap Between Student and TeacherYichen Zhu, Yi WangICCV 2021 · 95 citations
- ScaleKD: Distilling Scale-Aware Knowledge in Small Object DetectorYichen Zhu, Qiqi Zhou, Ning Liu, Zhiyuan Xu et al.CVPR 2023
