DETRDistill: A Universal Knowledge Distillation Framework for DETR-families
Jiahao Chang, Shuo Wang, Hai-Ming Xu, Zehui Chen, Chenhongyi Yang, Feng Zhao
Abstract
Transformer-based detectors (DETRs) are becoming popular for their simple framework, but the large model size and heavy time consumption hinder their deployment in the real world. While knowledge distillation (KD) can be an appealing technique to compress giant detectors into small ones for comparable detection performance and low inference cost. Since DETRs formulate object detection as a set prediction problem, existing KD methods designed for classic convolution-based detectors may not be directly applicable. In this paper, we propose DETRDistill, a novel knowledge distillation method dedicated to DETR-families. Specifically, we first design a Hungarian-matching logits distillation to encourage the student model to have the exact predictions as those of the teacher DETRs. Then, we propose a target-aware feature distillation to help the student model learn from the object-centric features of the teacher model. Finally, in order to improve the convergence rate of the student DETR, we introduce a query-prior assignment distillation to speed up the student model learning from well-trained queries and stable assignment of the teacher model. Extensive experimental results on the COCO dataset validate the effectiveness of our approach. Notably, DETRDistill consistently improves various DETRs by more than 2.0 mAP, even surpassing their teacher models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 90134168-af15-4640-9619-07e8a9075a4bCited by top-tier papers12
- SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal FusionMing Dai, Lingfeng Yang, Yihao Xu, Zhenhua Feng et al.NeurIPS 2024 · 67 citations
- Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation ModelsShenghao Fu, Junkai Yan, Qize Yang, Xihan Wei et al.NeurIPS 2024 · 24 citations
- KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points SamplingYu Wang, Xin Li, Shengzhao Weng, Gang Zhang et al.CVPR 2024 · 17 citations
- ASAG: Building Strong One-Decoder-Layer Sparse Detectors via Adaptive Sparse Anchor GenerationShenghao Fu, Junkai Yan, Yipeng Gao, Xiaohua Xie et al.ICCV 2023 · 8 citations
- AMap: Distilling Future Priors for Ahead-Aware Online HD Map ConstructionRuikai Li, Xinrun Li, Mengwei Xie, Hao Shan et al.CVPR 2026 · 8 citations
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
Related papers
- UniKD: Universal Knowledge Distillation for Mimicking Homogeneous or Heterogeneous Object DetectorsShanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu et al.ICCV 2023 · 7 citations
- Distilling Object Detectors with Feature RichnessZhixing Du, Rui Zhang, Ming Chang, Xishan Zhang et al.NeurIPS 2021 · 107 citations
- General Instance Distillation for Object DetectionXing Dai, Zeren Jiang, Zhao Wu, Yiping Bao et al.CVPR 2021
- CrossKD: Cross-Head Knowledge Distillation for Object DetectionJiabao Wang, Yuming Chen, Zhaohui Zheng, Xiang Li et al.CVPR 2024 · 93 citations
- ScaleKD: Distilling Scale-Aware Knowledge in Small Object DetectorYichen Zhu, Qiqi Zhou, Ning Liu, Zhiyuan Xu et al.CVPR 2023
