Semantics-Aware Dynamic Localization and Refinement for Referring Image Segmentation
Zhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen, Hengshuang Zhao, Philip H. S. Torr
Abstract
Referring image segmentation segments an image from a language expression. With the aim of producing high-quality masks, existing methods often adopt iterative learning approaches that rely on RNNs or stacked attention layers to refine vision-language features. Despite their complexity, RNN-based methods are subject to specific encoder choices, while attention-based methods offer limited gains. In this work, we introduce a simple yet effective alternative for progressively learning discriminative multi-modal features. The core idea of our approach is to leverage a continuously updated query as the representation of the target object and at each iteration, strengthen multi-modal features strongly correlated to the query while weakening less related ones. As the query is initialized by language features and successively updated by object features, our algorithm gradually shifts from being localization-centric to segmentation-centric. This strategy enables the incremental recovery of missing object parts and/or removal of extraneous parts through iteration. Compared to its counterparts, our method is more versatile—it can be plugged into prior arts straightforwardly and consistently bring improvements. Experimental results on the challenging datasets of RefCOCO, RefCOCO+, and G-Ref demonstrate its advantage with respect to the state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94e0edca-170f-4f42-b1a7-9992e441a528Cited by top-tier papers7
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image SegmentationZhenjie Mao, Yuhuan Yang, Chaofan Ma, Dongsheng Jiang et al.NeurIPS 2025 · 4 citations
- GenMask: Adapting DiT for Segmentation via Direct Mask GenerationYuhuan Yang, Xianwei Zhuang, Yuxuan Cai, Chaofan Ma et al.CVPR 2026 · 4 citations
- CoSTA: End-to-End Comprehensive Space-Time Entanglement for Spatio-Temporal Video GroundingYaoyuan Liang, Xiao Liang, Yansong Tang, Zhao Yang et al.AAAI 2024 · 3 citations
- Deep Instruction Tuning for Segment Anything ModelXiaorui Huang, Gen Luo, Chaoyang Zhu, Bo Tong et al.ACM MM 2024 · 3 citations
- Dynamic Prompting of Frozen Text-to-Image Diffusion Models for Panoptic Narrative GroundingHongyu Li, Tianrui Hui, Zihan Ding, Jing Zhang et al.ACM MM 2024 · 2 citations
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- Vision-Language Transformer and Query Generation for Referring SegmentationHenghui Ding, Chang Liu, Suchen Wang, Xudong JiangICCV 2021 · 359 citations
- UniT: Multimodal Multitask Learning with a Unified TransformerRonghang Hu, Amanpreet SinghICCV 2021 · 354 citations
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao et al.CVPR 2022 · 337 citations
Related papers
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Locate Then Segment: A Strong Pipeline for Referring Image SegmentationYa Jing, Tao Kong, Wei Wang, Liang Wang et al.CVPR 2021
- CARIS: Context-Aware Referring Image SegmentationSun'ao Liu, Yiheng Zhang, Zhaofan Qiu, Hongtao Xie et al.ACM MM 2023 · 34 citations
- Mask Grounding for Referring Image SegmentationYong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu et al.CVPR 2024
- LQMFormer: Language-Aware Query Mask Transformer for Referring Image SegmentationNisarg A. Shah, Vibashan VS, Vishal M. PatelCVPR 2024 · 11 citations
