Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETR
Feng Li, Ailing Zeng, Shilong Liu, Hao Zhang, Hongyang Li, Lei Zhang, Lionel M. Ni
Abstract
Recent DEtection TRansformer-based (DETR) models have obtained remarkable performance. Its success cannot be achieved without the re-introduction of multi-scale feature fusion in the encoder. However, the excessively increased tokens in multi-scale features, especially for about 75% of lowlevel features, are quite computationally inefficient, which hinders real applications of DETR models. In this paper, we present Lite DETR, a simple yet efficient end-to-end object detection framework that can effectively reduce the GFLOPs of the detection head by 60% while keeping 99% of the original performance. Specifically, we design an efficient encoder block to update high-level features (corresponding to smallresolution feature maps) and low-level features (corresponding to large-resolution feature maps) in an interleaved way. In addition, to better fuse cross-scale features, we develop a key-aware deformable attention to predict more reliable attention weights. Comprehensive experiments validate the effectiveness and efficiency of the proposed Lite DETR, and the efficient encoder strategy can generalize well across existing DETR-based models. The code will be available in https://github.com/IDEA-Research/Lite- DETR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2763b1d9-18ff-41db-bdc0-61318d0cc3a7Cited by top-tier papers15
- DETRs Beat YOLOs on Real-time Object DetectionYian Zhao, Wenyu Lv, Shangliang Xu, Jinman Wei et al.CVPR 2024 · 3,046 citations
- TAPTRv2: Attention-based Position Update Improves Tracking Any PointHongyang Li, Hao Zhang, Shilong Liu, Zhaoyang Zeng et al.NeurIPS 2024 · 22 citations
- DI-MaskDINO: A Joint Object Detection and Instance Segmentation ModelZhixiong Nan, Xianghong Li, Tao Xiang, Jifeng DaiNeurIPS 2024 · 15 citations
- AQ-DETR: Low-Bit Quantized Detection Transformer with Auxiliary QueriesRunqi Wang, Huixin Sun, Linlin Yang, Shaohui Lin et al.AAAI 2024 · 10 citations
- Referencing Where to Focus: Improving Visual Grounding with Referential QueryYabing Wang, Zhuotao Tian, Qingpei Guo, Zheng Qin et al.NeurIPS 2024 · 9 citations
Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
- Conditional DETR for Fast Training ConvergenceDepu Meng, Xiaokang Chen, Zejia Fan, Gang Zeng et al.ICCV 2021 · 974 citations
Related papers
- Sparse DETR: Efficient End-to-End Object Detection with Learnable SparsityByungseok Roh, Jaewoong Shin, Wuhyun Shin, Saehoon KimICLR 2022 · 256 citations
- CF-DETR: Coarse-to-Fine Transformers for End-to-End Object DetectionXipeng Cao, Peng Yuan, Bailan Feng, Kun NiuAAAI 2022 · 59 citations
- Mobile-Former: Bridging MobileNet and TransformerYinpeng Chen, Xiyang Dai, Dongdong Chen, Mengchen Liu et al.CVPR 2022 · 600 citations
- Deeply Tensor Compressed Transformers for End-to-End Object DetectionPeining Zhen, Ziyang Gao, Tianshu Hou, Yuan Cheng et al.AAAI 2022 · 19 citations
- Towards Efficient Use of Multi-Scale Features in Transformer-Based Object DetectorsGongjie Zhang, Zhipeng Luo, Zichen Tian, Jingyi Zhang et al.CVPR 2023
