CARIS: Context-Aware Referring Image Segmentation
Sun'ao Liu, Yiheng Zhang, Zhaofan Qiu, Hongtao Xie, Yongdong Zhang, Ting Yao
Abstract
Referring image segmentation aims to segment the target object described by a natural-language utterance. Recent approaches typically distinguish pixels by aligning pixel-wise visual features with linguistic features extracted from the referring description. Nevertheless, such a free-form description only specifies certain discriminative attributes of the target object or its relations to a limited number of objects, which fails to represent the rich visual context adequately. The stand-alone linguistic features are therefore unable to align with all visual concepts, resulting in inaccurate segmentation. In this paper, we propose to address this issue by incorporating rich visual context into linguistic features for sufficient vision-language alignment. Specifically, we present Context-Aware Referring Image Segmentation (CARIS), a novel architecture that enhances the contextual awareness of linguistic features via sequential vision-language attention and learnable prompts. Technically, CARIS develops a context-aware mask decoder with sequential bidirectional cross-modal attention to integrate the linguistic features with visual context, which are then aligned with pixel-wise visual features. Furthermore, two groups of learnable prompts are employed to delve into additional contextual information from the input image and facilitate the alignment with non-target pixels, respectively. Extensive experiments demonstrate that CARIS achieves new state-of-the-art performances on three public benchmarks. Code is available at https://github.com/lsa1997/CARIS.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers12
- RemoteSAM: Towards Segment Anything for Earth ObservationLiang Yao, Fan Liu, Delong Chen, Chuanyi Zhang et al.ACM MM 2025 · 28 citations
- Multi-task Visual Grounding with Coarse-to-Fine Consistency ConstraintsMing Dai, Jian Li, Jiedong Zhuang, Xian Zhang et al.AAAI 2025 · 23 citations
- Referencing Where to Focus: Improving Visual Grounding with Referential QueryYabing Wang, Zhuotao Tian, Qingpei Guo, Zheng Qin et al.NeurIPS 2024 · 9 citations
- 3D-GRES: Generalized 3D Referring Expression SegmentationChangli Wu, Yihang Liu, Jiayi Ji, Yiwei Ma et al.ACM MM 2024 · 7 citations
- MaskPrompt: Open-Vocabulary Affordance Segmentation with Object Shape Mask PromptsDongpan Chen, Dehui Kong, Jinghua Li, Baocai YinAAAI 2025 · 5 citations
Related papers
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Prompt-Driven Referring Image Segmentation with Instance ContrastingChao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang et al.CVPR 2024 · 20 citations
- Two-stage Visual Cues Enhancement Network for Referring Image SegmentationYang Jiao, Zequn Jie, Weixin Luo, Jingjing Chen et al.ACM MM 2021 · 24 citations
- Locate Then Segment: A Strong Pipeline for Referring Image SegmentationYa Jing, Tao Kong, Wei Wang, Liang Wang et al.CVPR 2021
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao et al.CVPR 2022 · 337 citations
