Structural Knowledge Distillation for Object Detection
Philip de Rijk, Lukas Schneider, Marius Cordts, Dariu Gavrila
Abstract
Knowledge Distillation (KD) is a well-known training paradigm in deep neural networks where knowledge acquired by a large teacher model is transferred to a small student. KD has proven to be an effective technique to significantly improve the student's performance for various tasks including object detection. As such, KD techniques mostly rely on guidance at the intermediate feature level, which is typically implemented by minimizing an p -norm distance between teacher and student activations during training. In this paper, we propose a replacement for the pixel-wise independent p -norm based on the structural similarity (SSIM) [28] . By taking into account additional contrast and structural cues, feature importance, correlation and spatial dependence in the feature space are considered in the loss formulation. Extensive experiments on MSCOCO [16] demonstrate the effectiveness of our method across different training schemes and architectures. Our method adds only little computational overhead, is straightforward to implement and at the same time it significantly outperforms the standard p -norms. Moreover, more complex state-of-the-art KD methods [13, 33] using attention-based sampling mechanisms are outperformed, including a +3.5 AP gain using a Faster R-CNN R-50 [21] compared to a vanilla model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li et al.CVPR 2024 · 93 citations
- DetKDS: Knowledge Distillation Search for Object DetectorsLujun Li, Yufan Bao, Peijie Dong, Chuanguang Yang et al.ICML 2024 · 35 citations
- FM-OV3D: Foundation Model-Based Cross-Modal Knowledge Blending for Open-Vocabulary 3D DetectionDongmei Zhang, Chang Li, Renrui Zhang, Shenghao Xie et al.AAAI 2024 · 24 citations
- How2Compress: Scalable and Efficient Edge Video Analytics via Adaptive Granular Video CompressionYuheng Wu, Thanh-Tung Nguyen, Lucas Liebe, Quang Tau et al.ACM MM 2025 · 1 citation
Builds on5
- RepPoints: Point Set Representation for Object DetectionZe Yang, Shaohui Liu, Han Hu, Liwei Wang et al.ICCV 2019 · 1,056 citations
- Improve Object Detection with Feature-based Knowledge Distillation: Towards Accurate and Efficient DetectorsLinfeng Zhang, Kaisheng MaICLR 2021 · 251 citations
- Distilling Object Detectors with Feature RichnessZhixing Du, Rui Zhang, Ming Chang, Xishan Zhang et al.NeurIPS 2021 · 107 citations
- Instance-Conditional Knowledge Distillation for Object DetectionZijian Kang, Peizhen Zhang, Xiangyu Zhang, Jian Sun et al.NeurIPS 2021 · 91 citations
- Distilling Object Detectors via Decoupled FeaturesJianyuan Guo, Kai Han, Yunhe Wang, Han Wu et al.CVPR 2021
Related papers
- Channel-wise Knowledge Distillation for Dense Prediction*Changyong Shu, Yifan Liu, Jianfei Gao, Zheng Yan et al.ICCV 2021 · 432 citations
- G-DetKD: Towards General Distillation Framework for Object Detectors via Contrastive and Semantic-guided Feature ImitationLewei Yao, Renjie Pi, Hang Xu, Wei Zhang et al.ICCV 2021 · 48 citations
- Knowledge Distillation from A Stronger TeacherTao Huang, Shan You, Fei Wang, Chen Qian et al.NeurIPS 2022 · 477 citations
- Similarity-Preserving Knowledge DistillationFrederick Tung, Greg MoriICCV 2019 · 1,214 citations
- Beyond the limitation of monocular 3D detector via knowledge distillationYiran Yang, Dongshuo Yin, Xuee Rong, Xian Sun et al.ICCV 2023 · 4 citations
