Adaptive Selection based Referring Image Segmentation
Pengfei Yue, Jianghang Lin, Shengchuan Zhang, Jie Hu, Yilin Lu, Hongwei Niu, Haixin Ding, Yan Zhang, Guannan Jiang, Liujuan Cao, Rongrong Ji
Abstract
Referring image segmentation (RIS) aims to segment a particular region based on a specific expression. Existing one-stage methods have explored various fusion strategies, yet they encounter two significant issues. Primarily, most methods rely on manually selected visual features from the visual encoder layers. Moreover, the direct fusion of word-level features into coarse aligned features disrupts the established vision-language alignment. In this paper, we introduce an innovative framework for RIS that seeks to overcome these challenges with adaptive alignment of vision and language features, termed the Adaptive Selection with Dual Alignment (ASDA). ASDA innovates in two aspects. Firstly, we design an Adaptive Feature Selection and Fusion (AFSF) module to dynamically select visual features focusing on different regions related to various descriptions. AFSF is equipped with scale-wise feature aggregator to provide hierarchically coarse features that preserve crucial low-level details. Secondly, a Word Guided Dual-Branch Aligner (WGDA) is leveraged to integrate coarse features with linguistic cues by word-guided attention, which effectively addresses the common issue of vision-language misalignment. Extensive experimental results demonstrate that our ASDA framework surpasses state-of-the-art methods on RefCOCO, RefCOCO+ and G-Ref benchmark.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 7e0d2ad1-48c1-4888-aaf2-3d5d6cfb7ac7Cited by top-tier papers4
- Exploring Semantic Consistency and Style Diversity for Domain Generalized Semantic SegmentationHongwei Niu, Linhuang Xie, Jianghang Lin, Shengchuan ZhangAAAI 2025 · 16 citations
- Generate Aligned Anomaly: Region-Guided Few-Shot Anomaly Image-Mask Pair Synthesis for Industrial InspectionYilin Lu, Jianghang Lin, Linhuang Xie, Kai Zhao et al.ACM MM 2025
- RIS-LAD: A Benchmark and Model for Referring Image Segmentation in Low-Altitude Drone ImageryKai Ye, YingShi Luan, Zhudi Chen, Guangyue Meng et al.AAAI 2026
- AnomalyPainter: Vision-Language-Diffusion Synergy for Realistic and Diverse Unseen Industrial Anomaly SynthesisZhangyu Lai, Yilin Lu, Xinyang Li, Jianghang Lin et al.AAAI 2026
Related papers
- Semantics-Aware Dynamic Localization and Refinement for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.AAAI 2023 · 35 citations
- AMLRIS: Alignment-aware Masked Learning for Referring Image SegmentationTongfei Chen, Shuo Yang, Yuguang Yang, Linlin Yang et al.ICLR 2026 · 1 citation
- Bottom-Up and Bidirectional Alignment for Referring Expression ComprehensionLiuwu Li, Yuqi Bu, Yi CaiACM MM 2021 · 11 citations
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen et al.CVPR 2022 · 319 citations
- Two-stage Visual Cues Enhancement Network for Referring Image SegmentationYang Jiao, Zequn Jie, Weixin Luo, Jingjing Chen et al.ACM MM 2021 · 24 citations
