ScanFormer: Referring Expression Comprehension by Iteratively Scanning
Wei Su, Peihan Miao, Huanzhang Dou, Xi Li
Abstract
Referring Expression Comprehension (REC) aims to localize the target objects specified by free-form natural language descriptions in images. While state-of-the-art methods achieve impressive performance, they perform a dense perception of images, which incorporates redundant visual regions unrelated to linguistic queries, leading to additional computational overhead. This inspires us to explore a question: can we eliminate linguistic-irrelevant redundant visual regions to improve the efficiency of the model? Existing relevant methods primarily focus on fundamental visual tasks, with limited exploration in vision-language fields. To address this, we propose a coarse-to-fine iterative perception framework, called ScanFormer. It can iteratively exploit the image scale pyramid to extract linguistic-relevant visual patches from top to bottom. In each iteration, irrelevant patches are discarded by our designed informativeness prediction. Furthermore, we propose a patch selection strategy for discarded patches to accelerate inference. Experiments on widely used datasets, namely Ref COCO, Ref COCO+, Ref COCO g, and ReferItGame, verify the effectiveness of our method, which can strike a balance between accuracy and efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 63c9a724-6830-4d9b-b9bb-28c2fd594cdaCited by top-tier papers11
- Multi-task Visual Grounding with Coarse-to-Fine Consistency ConstraintsMing Dai, Jian Li, Jiedong Zhuang, Xian Zhang et al.AAAI 2025 · 23 citations
- MaPPER: Multimodal Prior-guided Parameter Efficient Tuning for Referring Expression ComprehensionTing Liu, Zunnan Xu, Yue Hu, Liangtao Shi et al.EMNLP 2024 · 6 citations
- Look Around Before Locating: Considering Content and Structure Information for Visual GroundingShiyi Zheng, Peizhi Zhao, Zhilong Zheng, Peihang He et al.AAAI 2025 · 3 citations
- LaplacianFormer: Rethinking Linear Attention with Laplacian KernelZhe Feng, Sen Lian, Changwei Wang, Muyang Zhang et al.ICLR 2026 · 3 citations
- PropVG: End-To-End Proposal-Driven Visual Grounding with Multi-Granularity DiscriminationMing Dai, Wenxuan Cheng, Jiedong Zhuang, Jiang-jiang Liu et al.ICCV 2025 · 3 citations
Builds on18
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- Referring Expression Comprehension Using Language Adaptive InferenceWei Su, Peihan Miao, Huanzhang Dou, Yongjian Fu et al.AAAI 2023 · 34 citations
- Bottom-Up and Bidirectional Alignment for Referring Expression ComprehensionLiuwu Li, Yuqi Bu, Yi CaiACM MM 2021 · 11 citations
- A Real-Time Cross-Modality Correlation Filtering Method for Referring Expression ComprehensionYue Liao, Si Liu, Guanbin Li, Fei Wang et al.CVPR 2020
- Leveraging Debiased Cross-Modal Attention Maps and Code-Based Reasoning for Zero-Shot Referring Expression ComprehensionJuntao Chen, Wen Shen, Zhihua Wei, Lijun Sun et al.ICCV 2025 · 1 citation
- LUNA: Language as Continuing Anchors for Referring Expression ComprehensionYaoyuan Liang, Zhao Yang, Yansong Tang, Jiashuo Fan et al.ACM MM 2023 · 9 citations
