One-Stage Visual Grounding via Semantic-Aware Feature Filter
Jiabo Ye, Xin Lin, Liang He, Dingbang Li, Qin Chen
Abstract
Visual grounding has attracted much attention with the popularity of vision language. Existing one-stage methods are far ahead of two-stage methods in speed. However, these methods fuse the textual feature and visual feature map by simply concatenation, which ignores the textual semantics and limits these models' ability in cross-modal understanding. To overcome this weakness, we propose a semantic-aware framework that utilizes both queries' structured knowledge and context-sensitive representations to filter the visual feature maps to localize the referents more accurately. Our framework contains an entity filter, an attribute filter, and a location filter. These three filters filter the input visual feature map step by step according to each query's aspects respectively. A grounding module further regresses the bounding boxes to localize the referential object. Experiments on various commonly used datasets show that our framework achieves a real-time inference speed and outperforms all state-of-the-art methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 212ed661-fb4a-4303-b787-92986e04ef6fCited by top-tier papers5
- Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual GroundingJiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang et al.CVPR 2022 · 89 citations
- HERO: HiErarchical spatio-tempoRal reasOning with Contrastive Action Correspondence for End-to-End Video Object GroundingMengze Li, Tianbao Wang, Haoyu Zhang, Shengyu Zhang et al.ACM MM 2022 · 25 citations
- Visual Grounding with Multi-modal Conditional AdaptationRuilin Yao, Shengwu Xiong, Yichen Zhao, Yi RongACM MM 2024 · 19 citations
- PPMN: Pixel-Phrase Matching Network for One-Stage Panoptic Narrative GroundingZihan Ding, Zi-han Ding, Tianrui Hui, Junshi Huang et al.ACM MM 2022 · 12 citations
- Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression ComprehensionYaxian Wang, Henghui Ding, Shuting He, Xudong Jiang et al.AAAI 2025 · 9 citations
Related papers
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningLi Yang, Yan Xu, Chunfeng Yuan, Wei Liu et al.CVPR 2022 · 146 citations
- Look Around Before Locating: Considering Content and Structure Information for Visual GroundingShiyi Zheng, Peizhi Zhao, Zhilong Zheng, Peihang He et al.AAAI 2025 · 3 citations
- QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual GroundingShengxin Chen, Gen Luo, Yiyi Zhou, Xiaoshuai Sun et al.ACM MM 2024 · 6 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 270 citations
