Towards Efficient Use of Multi-Scale Features in Transformer-Based Object Detectors
Gongjie Zhang, Zhipeng Luo, Zichen Tian, Jingyi Zhang, Xiaoqin Zhang, Shijian Lu
Abstract
Multi-scale features have been proven highly effective for object detection but often come with huge and even prohibitive extra computation costs, especially for the recent Transformer-based detectors. In this paper, we propose Iterative Multi-scale Feature Aggregation (IMFA) -a generic paradigm that enables efficient use of multi-scale features in Transformer-based object detectors. The core idea is to exploit sparse multi-scale features from just a few crucial locations, and it is achieved with two novel designs. First, IMFA rearranges the Transformer encoderdecoder pipeline so that the encoded features can be iteratively updated based on the detection predictions. Second, IMFA sparsely samples scale-adaptive features for refined detection from just a few keypoint locations under the guidance of prior detection predictions. As a result, the sampled multi-scale features are sparse yet still highly beneficial for object detection. Extensive experiments show that the proposed IMFA boosts the performance of multiple Transformer-based object detectors significantly yet with only slight computational overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb16305a-f88a-400f-ac96-e9b2812c16e2Cited by top-tier papers9
- Less is More: Focus Attention for Efficient DETRDehua Zheng, Wenhui Dong, Hailin Hu, Xinghao Chen et al.ICCV 2023 · 128 citations
- Online Map Vectorization for Autonomous Driving: A Rasterization PerspectiveGongjie Zhang, Jiahao Lin, Shuang Wu, Yilin Song et al.NeurIPS 2023 · 78 citations
- DETR Does Not Need Multi-Scale or Locality DesignYutong Lin, Yuhui Yuan, Zheng Zhang, Chen Li et al.ICCV 2023 · 27 citations
- ASAG: Building Strong One-Decoder-Layer Sparse Detectors via Adaptive Sparse Anchor GenerationShenghao Fu, Junkai Yan, Yipeng Gao, Xiaohua Xie et al.ICCV 2023 · 8 citations
- OralXrays-9: Towards Hospital-Scale Panoramic X-ray Anomaly Detection via Personalized Multi-Object Query-Aware MiningBingzhi Chen, Sisi Fu, Xiaocheng Fang, Jieyi Cai et al.CVPR 2025
Builds on37
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
Related papers
- Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETRFeng Li, Ailing Zeng, Shilong Liu, Hao Zhang et al.CVPR 2023
- Sparse DETR: Efficient End-to-End Object Detection with Learnable SparsityByungseok Roh, Jaewoong Shin, Wuhyun Shin, Saehoon KimICLR 2022 · 256 citations
- Building Vision Transformers with Hierarchy Aware Feature AggregationYongjie Chen, Hongmin Liu, Haoran Yin, Bin FanICCV 2023 · 6 citations
- Scene Adaptive Sparse Transformer for Event-based Object DetectionYansong Peng, Hebei Li, Yueyi Zhang, Xiaoyan Sun et al.CVPR 2024 · 25 citations
- Integrally Migrating Pre-trained Transformer Encoder-decoders for Visual Object DetectionFeng Liu, Xiaosong Zhang, Zhiliang Peng, Zonghao Guo et al.ICCV 2023 · 30 citations
