KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points Sampling
Yu Wang, Xin Li, Shengzhao Weng, Gang Zhang, Haixiao Yue, Haocheng Feng, Junyu Han, Errui Ding
Abstract
DETR is a novel end-to-end transformer architecture object detector, which significantly outperforms classic detectors when scaling up. In this paper, we focus on the compression of DETR with knowledge distillation. While knowledge distillation has been well-studied in classic detectors, there is a lack of researches on how to make it work effectively on DETR. We first provide experimental and theoretical analysis to point out that the main challenge in DETR distillation is the lack of consistent distillation points. Distillation points refer to the corresponding inputs of the predictions for student to mimic, which have different formulations in CNN detector and DETR, and reliable distillation requires sufficient distillation points which are consistent between teacher and student. Based on this observation, we propose the first general knowledge distillation paradigm for DETR (KD-DETR) with consistent distillation points sampling, for both homogeneous and heterogeneous distillation. Specifically, we decouple detection and distillation tasks by introducing a set of specialized object queries to construct distillation points for DETR. We further propose a general-to-specific distillation points sampling strategy to explore the extensibility of KD-DETR. Extensive experiments validate the effectiveness and generalization of KD-DETR. For both single-scale DAB-DETR and multis-scale Deformable DETR and DINO, KD-DETR boost the performance of student model with improvements of 2.6% - 5.2%. We further extend KD-DETR to heterogeneous distillation, and achieves 2.1 % improvement by distilling the knowledge from DINO to Faster R-CNN with ResNet-50, which is comparable with homogeneous distillation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5055b32f-c39e-4c5c-8159-e20bb260649dCited by top-tier papers5
- GCD: Advancing Vision-Language Models for Incremental Object Detection via Global Alignment and Correspondence DistillationXu Wang, Zilei Wang, Zihan LinAAAI 2025 · 4 citations
- Dual Domain Control via Active Learning for Remote Sensing Domain Incremental Object DetectionJiachen Sun, De Cheng, Xi Yang, Nannan WangICCV 2025 · 2 citations
- D-FINE: Redefine Regression Task of DETRs as Fine-grained Distribution RefinementYansong Peng, Hebei Li, Peixi Wu, Yueyi Zhang et al.ICLR 2025
- Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and CompressionNazia Tasnim, Shrimai Prabhumoye, Bryan A. PlummerCVPR 2026
- Symbiosis-Inspired Knowledge Distillation for Incremental Object DetectionMingyue Zeng, De Cheng, Zhipeng Xu, Huaijie Wang et al.ICML 2026
Builds on22
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen et al.NeurIPS 2020 · 2,118 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
Related papers
- DETRDistill: A Universal Knowledge Distillation Framework for DETR-familiesJiahao Chang, Shuo Wang, Hai-Ming Xu, Zehui Chen et al.ICCV 2023 · 53 citations
- UniKD: Universal Knowledge Distillation for Mimicking Homogeneous or Heterogeneous Object DetectorsShanshan Lao, Guanglu Song, Boxiao Liu, Yu Liu et al.ICCV 2023 · 7 citations
- G-DetKD: Towards General Distillation Framework for Object Detectors via Contrastive and Semantic-guided Feature ImitationLewei Yao, Renjie Pi, Hang Xu, Wei Zhang et al.ICCV 2021 · 48 citations
- PKD: General Distillation Framework for Object Detectors via Pearson Correlation CoefficientWeihan Cao, Yifan Zhang, Jianfei Gao, Anda Cheng et al.NeurIPS 2022 · 147 citations
- DFD: Distilling the Feature Disparity Differently for DetectorsKang Liu, Yingyi Zhang, Jingyun Zhang, Jinmin Li et al.ICML 2024
