SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
Zhenjie Mao, Yuhuan Yang, Chaofan Ma, Dongsheng Jiang, Jiangchao Yao, Ya Zhang, Yanfeng Wang
Abstract
Referring Image Segmentation (RIS) aims to segment the target object in an image given a natural language expression. While recent methods leverage pre-trained vision backbones and more training corpus to achieve impressive results, they predominantly focus on simple expressions-short, clear noun phrases like "red car" or "left girl". This simplification often reduces RIS to a key word/concept matching problem, limiting the model's ability to handle referential ambiguity in expressions. In this work, we identify two challenging real-world scenarios: object-distracting expressions, which involve multiple entities with contextual cues, and category-implicit expressions, where the object class is not explicitly stated. To address the challenges, we propose a novel framework, SaFiRe, which mimics the human two-phase cognitive process-first forming a global understanding, then refining it through detail-oriented inspection. This is naturally supported by Mamba's scan-then-update property, which aligns with our phased design and enables efficient multi-cycle refinement with linear complexity. We further introduce aRefCOCO, a new benchmark designed to evaluate RIS models under ambiguous referring expressions. Extensive experiments on both standard and proposed datasets demonstrate the superiority of SaFiRe over state-of-the-art baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- GenMask: Adapting DiT for Segmentation via Direct Mask GenerationYuhuan Yang, Xianwei Zhuang, Yuxuan Cai, Chaofan Ma et al.CVPR 2026 · 4 citations
- Reason, Then Re-reason: Cross-view Revisiting Improves Spatial ReasoningChaofan Ma, Zhenjie Mao, Yuhuan Yang, Fanqin Zeng et al.ICML 2026 · 1 citation
Builds on43
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
Related papers
- Mask Grounding for Referring Image SegmentationYong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu et al.CVPR 2024
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation Through Loopback SynergyMing Dai, Wenxuan Cheng, Jiang-Jiang Liu, Sen Yang et al.ICCV 2025 · 5 citations
- Bottom-Up Shift and Reasoning for Referring Image SegmentationSibei Yang, Meng Xia, Guanbin Li, Hong-Yu Zhou et al.CVPR 2021
- Locate Then Segment: A Strong Pipeline for Referring Image SegmentationYa Jing, Tao Kong, Wei Wang, Liang Wang et al.CVPR 2021
- Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationWenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo et al.CVPR 2024 · 5 citations
