Less is More: Focus Attention for Efficient DETR
Dehua Zheng, Wenhui Dong, Hailin Hu, Xinghao Chen, Yunhe Wang
Abstract
DETR-like models have significantly boosted the performance of detectors and even outperformed classical convolutional models. However, all tokens are treated equally without discrimination brings a redundant computational burden in the traditional encoder structure. The recent sparsification strategies exploit a subset of informative tokens to reduce attention complexity maintaining performance through the sparse encoder. But these methods tend to rely on unreliable model statistics. Moreover, simply reducing the token population hinders the detection performance to a large extent, limiting the application of these sparse models. We propose Focus-DETR, which focuses attention on more informative tokens for a better trade-off between computation efficiency and model accuracy. Specifically, we reconstruct the encoder with dual attention, which includes a token scoring mechanism that considers both localization and category semantic information of the objects from multi-scale feature maps. We efficiently abandon the background queries and enhance the semantic interaction of the fine-grained object queries based on the scores. Compared with the state-of-the-art sparse DETR-like detectors under the same setting, our Focus-DETR gets comparable complexity while achieving 50.4AP (+2.2) on COCO. The code is available at https://github.com/huawei-noah/noah-research/tree/master/Focus-DETR and https://gitee.com/mindspore/models/tree/master/research/cv/Focus-DETR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6353d7ba-d676-44fc-afda-993701f755f2Cited by top-tier papers19
- SLAB: Efficient Transformers with Simplified Linear Attention and Progressive Re-parameterized Batch NormalizationJialong Guo, Xinghao Chen, Yehui Tang, Yunhe WangICML 2024 · 40 citations
- Dynamic Dictionary Learning for Remote Sensing Image SegmentationXuechao Zou, Yue Li, Shun Zhang, Kai Li et al.ICCV 2025 · 15 citations
- DI-MaskDINO: A Joint Object Detection and Instance Segmentation ModelZhixiong Nan, Xianghong Li, Tao Xiang, Jifeng DaiNeurIPS 2024 · 15 citations
- ERQ: Error Reduction for Post-Training Quantization of Vision TransformersYunshan Zhong, Jiawei Hu, You Huang, Yuxin Zhang et al.ICML 2024 · 14 citations
- DON'T NEED RETRAINING: A Mixture of DETR and Vision Foundation Models for Cross-Domain Few-Shot Object DetectionChanghan Liu, Xunzhi Xiang, Zixuan Duan, Wenbin Li et al.NeurIPS 2025 · 8 citations
Builds on22
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
Related papers
- Sparse DETR: Efficient End-to-End Object Detection with Learnable SparsityByungseok Roh, Jaewoong Shin, Wuhyun Shin, Saehoon KimICLR 2022 · 256 citations
- Not All Tokens Matter All The Time: Dynamic Token Aggregation Towards Efficient Detection TransformersJiacheng Cheng, Xiwen Yao, Xiang Yuan, Junwei HanICML 2025
- Salience DETR: Enhancing Detection Transformer with Hierarchical Salience Filtering RefinementXiuquan Hou, Meiqin Liu, Senlin Zhang, Ping Wei et al.CVPR 2024
- Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETRFeng Li, Ailing Zeng, Shilong Liu, Hao Zhang et al.CVPR 2023
- Dynamic Focus-aware Positional Queries for Semantic SegmentationHaoyu He, Jianfei Cai, Zizheng Pan, Jing Liu et al.CVPR 2023
