DETR Does Not Need Multi-Scale or Locality Design
Yutong Lin, Yuhui Yuan, Zheng Zhang, Chen Li, Nanning Zheng, Han Hu
Abstract
This paper presents an improved DETR detector that maintains a "plain" nature: using a single-scale feature map and global cross-attention calculations without specific locality constraints, in contrast to previous leading DETR-based detectors that reintroduce architectural inductive biases of multi-scale and locality into the decoder. We show that two simple technologies are surprisingly effective within a plain design to compensate for the lack of multiscale feature maps and locality constraints. The first is a box-to-pixel relative position bias (BoxRPB) term added to the cross-attention formulation, which well guides each query to attend to the corresponding object region while also providing encoding flexibility. The second is masked image modeling (MIM)-based backbone pre-training which helps learn representation with fine-grained localization ability and proves crucial for remedying dependencies on the multi-scale feature maps. By incorporating these technologies and recent advancements in training and problem formation, the improved "plain" DETR showed exceptional improvements over the original DETR detector. By leveraging the Object365 dataset for pre-training, it achieved 63.9 mAP accuracy using a Swin-L backbone, which is highly competitive with state-of-the-art detectors which all heavily rely on multi-scale feature maps and region-based feature extraction. Code will be available at https:// github.com/ impiga/ Plain-DETR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3bc2c7f-ad6c-4b14-b022-add640be2054Cited by top-tier papers9
- Rank-DETR for High Quality Object DetectionYifan Pu, Weicong Liang, Yiduo Hao, Yuhui Yuan et al.NeurIPS 2023 · 138 citations
- Hybrid Proposal Refiner: Revisiting DETR Series from the Faster R-CNN PerspectiveJinjing Zhao, Fangyun Wei, Chang XuCVPR 2024 · 14 citations
- Prediction-Feedback DETR for Temporal Action DetectionJihwan Kim, Miso Lee, Cheol-Ho Cho, Jihyun Lee et al.AAAI 2025 · 8 citations
- Revisiting [CLS] and Patch Token Interaction in Vision TransformersAlexis Marouani, Oriane Siméoni, Hervé Jégou, Piotr Bojanowski et al.ICLR 2026 · 6 citations
- DAMap: Distance-Aware MapNet for High Quality HD Map ConstructionJinpeng Dong, Chen Li, Yutong Lin, Jingwen Fu et al.ICCV 2025 · 1 citation
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
Related papers
- UP-DETR: Unsupervised Pre-Training for Object Detection With TransformersZhigang Dai, Bolun Cai, Yugeng Lin, Junying ChenCVPR 2021
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang et al.ICLR 2023 · 753 citations
- WB-DETR: Transformer-Based Detector without BackboneFanfan Liu, Haoran Wei, Wenzhe Zhao, Guozhen Li et al.ICCV 2021 · 45 citations
- Siamese DETRZeren Chen, Gengshi Huang, Wei Li, Jianing Teng et al.CVPR 2023
- CP-DETR: Concept Prompt Guide DETR Toward Stronger Universal Object DetectionQibo Chen, Weizhong Jin, Jianyue Ge, Mengdi Liu et al.AAAI 2025 · 3 citations
