DAC-DETR: Divide the Attention Layers and Conquer
Zhengdong Hu, Yifan Sun, Jingdong Wang, Yi Yang
摘要
This paper reveals a characteristic of DEtection Transformer (DETR) that negatively impacts its training efficacy, i.e. , the cross-attention and self-attention layers in DETR decoder have opposing impacts on the object queries (though both impacts are important). Specifically, we observe the cross-attention tends to gather multiple queries around the same object, while the self-attention disperses these queries far away. To improve the training efficacy, we propose a Divide-And-Conquer DETR (DAC-DETR) that separates out the cross-attention to avoid these competing objectives. During training, DAC-DETR employs an auxiliary decoder that focuses on learning the cross-attention layers. The auxiliary decoder, while sharing all the other parameters, has NO self-attention layers and employs one-to-many label assignment to improve the gathering effect. Experiments show that DAC-DETR brings remarkable improvement over popular DETRs. For example, under the 12 epochs training scheme on MS-COCO, DAC-DETR improves Deformable DETR (ResNet-50) by +3.4AP and achieves 50.9 (ResNet-50) / 58.1 AP (Swin-Large) based on some popular methods ( i.e. , DINO and an IoU-related loss). Our code will be made available at https://github.com/huzhengdongcs/DAC-DETR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath 等ICLR 2026 · 被引用 1,103 次
- MS-DETR: Efficient DETR Training with Mixed SupervisionChuyang Zhao, Yifan Sun, Wenhao Wang, Qiang Chen 等CVPR 2024 · 被引用 51 次
- Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation ModelsShenghao Fu, Junkai Yan, Qize Yang, Xihan Wei 等NeurIPS 2024 · 被引用 24 次
- RETR: Multi-View Radar Detection Transformer for Indoor PerceptionRyoma Yataka, Adriano Cardace, Perry Wang, Petros Boufounos 等NeurIPS 2024 · 被引用 21 次
- DI-MaskDINO: A Joint Object Detection and Instance Segmentation ModelZhixiong Nan, Xianghong Li, Tao Xiang, Jifeng DaiNeurIPS 2024 · 被引用 15 次
它引用的顶会 Paper22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 被引用 6,042 次
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang 等ICLR 2022 · 被引用 1,218 次
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng 等ICCV 2021 · 被引用 974 次
相关 Paper
- DETRs with Collaborative Hybrid Assignments TrainingZhuofan Zong, Guanglu Song, Yu LiuICCV 2023 · 被引用 594 次
- Group DETR: Fast DETR Training with Group-Wise One-to-Many AssignmentQiang Chen, Xiaokang Chen, Jian Wang, Shan Zhang 等ICCV 2023 · 被引用 231 次
- DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object DetectionGuiping Cao, Xiangyuan Lan, Wenjian Huang, Jianguo Zhang 等ACM MM 2025
- DESTR: Object Detection with Split TransformerLiqiang He, Sinisa TodorovicCVPR 2022 · 被引用 63 次
- Rank-DETR for High Quality Object DetectionYifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan 等NeurIPS 2023 · 被引用 138 次
