Dynamic DETR: End-to-End Object Detection with Dynamic Attention
Xiyang Dai, Yinpeng Chen, Jianwei Yang, Pengchuan Zhang, Lu Yuan, Lei Zhang
Abstract
In this paper, we present a novel Dynamic DETR (Detection with Transformers) approach by introducing dynamic attentions into both the encoder and decoder stages of DETR to break its two limitations on small feature resolution and slow training convergence. To address the first limitation, which is due to the quadratic computational complexity of the self-attention module in Transformer encoders, we propose a dynamic encoder to approximate the Transformer encoder's attention mechanism using a convolution-based dynamic encoder with various attention types. Such an encoder can dynamically adjust attentions based on multiple factors such as scale importance, spatial importance, and representation (i.e., feature dimension) importance. To mitigate the second limitation of learning difficulty, we introduce a dynamic decoder by replacing the cross-attention module with a ROI-based dynamic attention in the Transformer decoder. Such a decoder effectively assists Transformers to focus on region of interests from a coarse-to-fine manner and dramatically lowers the learning difficulty, leading to a much faster convergence with fewer training epochs. We conduct a series of experiments to demonstrate our advantages. Our Dynamic DETR significantly reduces the training epochs (by 14×), yet results in a much better performance (by 3.6 on mAP). Meanwhile, in the standard 1× setup with ResNet-50 backbone, we archive a new state-of-the-art performance that further proves the learning effectiveness of the proposed approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers79
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- DN-DETR: Accelerate DETR Training by Introducing Query DeNoisingFeng Li, Hao Zhang, Shilong Liu, Jian Guo et al.CVPR 2022 · 879 citations
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- Retinexformer: One-stage Retinex-based Transformer for Low-light Image EnhancementYuanhao Cai, Hao Bian, Jing Lin, Haoqian Wang et al.ICCV 2023 · 615 citations
- Mask-guided Spectral-wise Transformer for Efficient Hyperspectral Image ReconstructionYuanhao Cai, Jing Lin, Xiaowan Hu, Haoqian Wang et al.CVPR 2022 · 310 citations
Builds on9
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- ConvBERT: Improving BERT with Span-based Dynamic ConvolutionZihang Jiang, Weihao Yu, Daquan Zhou, Yunpeng Chen et al.NeurIPS 2020 · 220 citations
- RepPoints v2: Verification Meets Regression for Object DetectionYihong Chen, Zheng Zhang, Yue Cao, Liwei Wang et al.NeurIPS 2020 · 134 citations
- Dynamic Convolution: Attention Over Convolution KernelsYinpeng Chen, Xiyang Dai, Mengchen Liu, Dongdong Chen et al.CVPR 2020
Related papers
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- CF-DETR: Coarse-to-Fine Transformers for End-to-End Object DetectionXipeng Cao, Peng Yuan, Bailan Feng, Kun NiuAAAI 2022 · 59 citations
- Not All Tokens Matter All The Time: Dynamic Token Aggregation Towards Efficient Detection TransformersJiacheng Cheng, Xiwen Yao, Xiang Yuan, Junwei HanICML 2025
- Recurrent Glimpse-based Decoder for Detection with TransformerZhe Chen, Jing Zhang, Dacheng TaoCVPR 2022 · 37 citations
- DAC-DETR: Divide the Attention Layers and ConquerZhengdong Hu, Yifan Sun, Jingdong Wang, Yi YangNeurIPS 2023 · 53 citations
