Decoupled DETR: Spatially Disentangling Localization and Classification for Improved End-to-End Object Detection
Manyuan Zhang, Guanglu Song, Yu Liu, Hongsheng Li
Abstract
The introduction of DETR represents a new paradigm for object detection. However, its decoder conducts classification and box localization using shared queries and cross-attention layers, leading to suboptimal results. We observe that different regions of interest in the visual feature map are suitable for performing query classification and box localization tasks, even for the same object. Salient regions provide vital information for classification, while the boundaries around them are more favorable for box regression. Unfortunately, such spatial misalignment between these two tasks greatly hinders DETR’s training. Therefore, in this work, we focus on decoupling localization and classification tasks in DETR. To achieve this, we introduce a new design scheme called spatially decoupled DETR (SD-DETR), which includes a task-aware query generation module and a disentangled feature learning process. We elaborately design the task-aware query initialization process and divide the cross-attention block in the decoder to allow the task-aware queries to match different visual regions. Meanwhile, we also observe that the prediction misalignment problem for high classification confidence and precise localization exists, so we propose an alignment loss to further guide the spatially decoupled DETR training. Through extensive experiments, we demonstrate that our approach achieves a significant improvement in MSCOCO datasets compared to previous work. For instance, we improve the performance of Conditional DETR by 4.5 AP. By spatially disentangling the two tasks, our method overcomes the misalignment problem and greatly improves the performance of DETR for object detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37dfd632-be07-42af-b6a3-e4915666eaa1Cited by top-tier papers5
- DSD-DA: Distillation-based Source Debiasing for Domain Adaptive Object DetectionYongchao Feng, Shiwei Li, Yingjie Gao, Ziyue Huang et al.ICML 2024 · 11 citations
- IDseq: Decoupled and Sequentially Detecting and Grounding Multi-Modal Media ManipulationRunxin Liu, Tian Xie, Jiaming Li, Lingyun Yu et al.AAAI 2025 · 2 citations
- DAMap: Distance-Aware MapNet for High Quality HD Map ConstructionJinpeng Dong, Chen Li, Yutong Lin, Jingwen Fu et al.ICCV 2025 · 1 citation
- Event-Equalized Dense Video CaptioningKangyi Wu, Pengna Li, Jingwen Fu, Yizhe Li et al.CVPR 2025
- Mr. DETR: Instructive Multi-Route Training for Detection TransformersChang-Bin Zhang, Yujie Zhong, Kai HanCVPR 2025
Builds on20
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- Generalized Focal Loss: Learning Qualified and Distributed Bounding Boxes for Dense Object DetectionXiang Li, Wenhai Wang, Lijun Wu, Shuo Chen et al.NeurIPS 2020 · 2,118 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- TOOD: Task-aligned One-stage Object DetectionChengjian Feng, Yujie Zhong, Yu Gao, Matthew R. Scott et al.ICCV 2021 · 1,191 citations
Related papers
- Revisiting the Sibling Head in Object DetectorGuanglu Song, Yu Liu, Xiaogang WangCVPR 2020
- DESTR: Object Detection with Split TransformerLiqiang He, Sinisa TodorovicCVPR 2022 · 63 citations
- Dual DETRs for Multi-Label Temporal Action DetectionYuhan Zhu, Guozhen Zhang, Jing Tan, Gangshan Wu et al.CVPR 2024 · 25 citations
- Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering RefinementXiuquan Hou, Meiqin Liu, Senlin Zhang, Ping Wei et al.CVPR 2024
- DAC-DETR: Divide the Attention Layers and ConquerZhengdong Hu, Yifan Sun, Jingdong Wang, Yi YangNeurIPS 2023 · 53 citations
