Sparse DETR: Efficient End-to-End Object Detection with Learnable Sparsity
Byungseok Roh, Jaewoong Shin, Wuhyun Shin, Saehoon Kim
Abstract
DETR is the first end-to-end object detector using a transformer encoder-decoder architecture and demonstrates competitive performance but low computational efficiency on high resolution feature maps. The subsequent work, Deformable DETR, enhances the efficiency of DETR by replacing dense attention with deformable attention, which achieves 10× faster convergence and improved performance. Deformable DETR uses the multiscale feature to ameliorate performance, however, the number of encoder tokens increases by 20× compared to DETR, and the computation cost of the encoder attention remains a bottleneck. In our preliminary experiment, we observe that the detection performance hardly deteriorates even if only a part of the encoder token is updated. Inspired by this observation, we propose Sparse DETR that selectively updates only the tokens expected to be referenced by the decoder, thus help the model effectively detect objects. In addition, we show that applying an auxiliary detection loss on the selected tokens in the encoder improves the performance while minimizing computational overhead. We validate that Sparse DETR achieves better performance than Deformable DETR even with only 10% encoder tokens on the COCO dataset. Albeit only the encoder tokens are sparsified, the total computation cost decreases by 38% and the frames per second (FPS) increases by 42% compared to Deformable DETR. Code is available at https://github.com/kakaobrain/sparse-detr .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers46
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu et al.NeurIPS 2022 · 742 citations
- CRN: Camera Radar Net for Accurate, Robust, Efficient 3D PerceptionYoungseok Kim, Juyeb Shin, Sanmin Kim, In-Jae Lee et al.ICCV 2023 · 134 citations
- Less is More: Focus Attention for Efficient DETRDehua Zheng, Wenhui Dong, Hailin Hu, Xinghao Chen et al.ICCV 2023 · 128 citations
- ClusterFomer: Clustering As A Universal Visual LearnerJames Liang, Yiming Cui, Qifan Wang, Tong Geng et al.NeurIPS 2023 · 63 citations
Builds on8
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision TransformersBowen Pan, Rameswar Panda, Yifan Jiang, Zhangyang Wang et al.NeurIPS 2021 · 209 citations
Related papers
- Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETRFeng Li, Ailing Zeng, Shilong Liu, Hao Zhang et al.CVPR 2023
- Not All Tokens Matter All The Time: Dynamic Token Aggregation Towards Efficient Detection TransformersJiacheng Cheng, Xiwen Yao, Xiang Yuan, Junwei HanICML 2025
- CF-DETR: Coarse-to-Fine Transformers for End-to-End Object DetectionXipeng Cao, Peng Yuan, Bailan Feng, Kun NiuAAAI 2022 · 59 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
- DESTR: Object Detection with Split TransformerLiqiang He, Sinisa TodorovicCVPR 2022 · 63 citations
