Iterative Shrinking for Referring Expression Grounding Using Deep Reinforcement Learning
Mingjie Sun, Jimin Xiao, Eng Gee Lim
Abstract
In this paper, we are tackling the proposal-free referring expression grounding task, aiming at localizing the target object according to a query sentence, without relying on off-the-shelf object proposals. Existing proposal-free methods employ a query-image matching branch to select the highest-score point in the image feature map as the target box center, with its width and height predicted by another branch. Such methods, however, fail to utilize the contextual relation between the target and reference objects, and lack interpretability on its reasoning procedure. To solve these problems, we propose an iterative shrinking mechanism to localize the target, where the shrinking direction is decided by a reinforcement learning agent, with all contents within the current image patch comprehensively considered. Besides, the sequential shrinking processes enable to demonstrate the reasoning about how to iteratively find the target. Experiments show that the proposed method boosts the accuracy by 4.32% against the previous state-of-theart (SOTA) method on the RefCOCOg dataset, where query sentences are long and complex with many targets referred by other reference objects.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual GroundingJiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang et al.CVPR 2022 · 89 citations
- Parallel Vertex Diffusion for Unified Visual GroundingZesen Cheng, Kehan Li, Peng Jin, Siheng Li et al.AAAI 2024 · 41 citations
- Multi-Modal Dynamic Graph Transformer for Visual GroundingSijia Chen, Baochun LiCVPR 2022 · 27 citations
- Rethinking Two-Stage Referring Expression Comprehension: A Novel Grounding and Segmentation Method Modulated by PointPeizhi Zhao, Shiyi Zheng, Wenye Zhao, Dongsheng Xu et al.AAAI 2024 · 11 citations
- Correspondence Matters for Video Referring Expression ComprehensionMeng Cao, Ji Jiang, Long Chen, Yuexian ZouACM MM 2022 · 10 citations
Builds on24
- Visual Semantic Reasoning for Image-Text MatchingKunpeng Li, Yulun Zhang, Kai Li, Yuanyuan Li et al.ICCV 2019 · 598 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 317 citations
- Dynamic Graph Attention for Referring Expression ComprehensionSibei Yang, Guanbin Li, Yizhou YuICCV 2019 · 251 citations
- Zero-Shot Grounding of Objects From Natural Language QueriesArka Sadhu, Kan Chen, Ram NevatiaICCV 2019 · 176 citations
Related papers
- Adaptive Reconstruction Network for Weakly Supervised Referring Expression GroundingXuejing Liu, Liang Li, Shuhui Wang, Zheng-Jun Zha et al.ICCV 2019 · 93 citations
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai et al.AAAI 2026
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 270 citations
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought ReasoningQing Jiang, Xingyu Chen, Zhaoyang Zeng, Junzhi Yu et al.ICLR 2026 · 25 citations
- Towards Further Comprehension on Referring Expression with RationaleRengang Li, Baoyu Fan, Xiaochuan Li, Runze Zhang et al.ACM MM 2022 · 2 citations
