DFD: Distilling the Feature Disparity Differently for Detectors
Kang Liu, Yingyi Zhang, Jingyun Zhang, Jinmin Li, Jun Wang, Shaoming Wang, Chun Yuan, Rizen Guo
Abstract
Knowledge distillation is a widely adopted model compression technique that has been successfully applied to object detection. In feature distillation, it is common practice for the student model to imitate the feature responses of the teacher model, with the underlying objective of improving its own abilities by reducing the disparity with the teacher. However, it is crucial to recognize that the disparities between the student and teacher are inconsistent, highlighting their varying abilities. In this paper, we explore the inconsistency in the disparity between teacher and student feature maps and analyze their impact on the efficiency of the distillation. We find that regions with varying degrees of difference should be treated separately, with different distillation constraints applied accordingly. We introduce our distillation method called Disparity Feature Distillation(DFD). The core idea behind DFD is to apply different treatments to regions with varying learning difficulties, simultaneously incorporating leniency and strictness. It enables the student to better assimilate the teacher's knowledge. Through extensive experiments, we demonstrate the effectiveness of our proposed DFD in achieving significant improvements. For instance, when applied to detectors based on ResNet50 such as Reti-naNet, FasterRCNN, and RepPoints, our method enhances their mAP from 37.4%, 38.4%, 38.6% to 41.7%, 42.4%, 42.7%, respectively. Our approach also demonstrates substantial improvements on YOLO and ViT-based models. The code is available at https://github.com/luckin99/DFD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 449f067c-8cc0-45fd-a95c-c5ceae6fa50bBuilds on12
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- RepPoints: Point Set Representation for Object DetectionZe Yang, Shaohui Liu, Han Hu, Liwei Wang et al.ICCV 2019 · 1,056 citations
- Decoupled Knowledge DistillationBorui Zhao, Quan Cui, Renjie Song, Yiyu Qiu et al.CVPR 2022 · 835 citations
- Focal and Global Knowledge Distillation for DetectorsZhendong Yang, Zhe Li, Xiaohu Jiang, Yuan Gong et al.CVPR 2022 · 325 citations
- Cross-Image Relational Knowledge Distillation for Semantic SegmentationChuanguang Yang, Helong Zhou, Zhulin An, Xue Jiang et al.CVPR 2022 · 228 citations
Related papers
- G-DetKD: Towards General Distillation Framework for Object Detectors via Contrastive and Semantic-guided Feature ImitationLewei Yao, Renjie Pi, Hang Xu, Wei Zhang et al.ICCV 2021 · 48 citations
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li et al.CVPR 2024 · 93 citations
- PKD: General Distillation Framework for Object Detectors via Pearson Correlation CoefficientWeihan Cao, Yifan Zhang, Jianfei Gao, Anda Cheng et al.NeurIPS 2022 · 147 citations
- Knowledge Distillation for Object Detection via Rank Mimicking and Prediction-Guided Feature ImitationGang Li, Xiang Li, Yujie Wang, Shanshan Zhang et al.AAAI 2022 · 105 citations
- Localization Distillation for Dense Object DetectionZhaohui Zheng, Rongguang Ye, Ping Wang, Dongwei Ren et al.CVPR 2022 · 177 citations
