Adaptive Selection based Referring Image Segmentation
Pengfei Yue, Jianghang Lin, Shengchuan Zhang, Jie Hu, Yilin Lu, Hongwei Niu, Haixin Ding, Yan Zhang, Guannan Jiang, Liujuan Cao, Rongrong Ji
摘要
Referring image segmentation (RIS) aims to segment a particular region based on a specific expression. Existing one-stage methods have explored various fusion strategies, yet they encounter two significant issues. Primarily, most methods rely on manually selected visual features from the visual encoder layers. Moreover, the direct fusion of word-level features into coarse aligned features disrupts the established vision-language alignment. In this paper, we introduce an innovative framework for RIS that seeks to overcome these challenges with adaptive alignment of vision and language features, termed the Adaptive Selection with Dual Alignment (ASDA). ASDA innovates in two aspects. Firstly, we design an Adaptive Feature Selection and Fusion (AFSF) module to dynamically select visual features focusing on different regions related to various descriptions. AFSF is equipped with scale-wise feature aggregator to provide hierarchically coarse features that preserve crucial low-level details. Secondly, a Word Guided Dual-Branch Aligner (WGDA) is leveraged to integrate coarse features with linguistic cues by word-guided attention, which effectively addresses the common issue of vision-language misalignment. Extensive experimental results demonstrate that our ASDA framework surpasses state-of-the-art methods on RefCOCO, RefCOCO+ and G-Ref benchmark.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- Exploring Semantic Consistency and Style Diversity for Domain Generalized Semantic SegmentationHongwei Niu, Linhuang Xie, Jianghang Lin, Shengchuan ZhangAAAI 2025 · 被引用 16 次
- Generate Aligned Anomaly: Region-Guided Few-Shot Anomaly Image-Mask Pair Synthesis for Industrial InspectionYilin Lu, Jianghang Lin, Linhuang Xie, Kai Zhao 等ACM MM 2025
- RIS-LAD: A Benchmark and Model for Referring Image Segmentation in Low-Altitude Drone ImageryKai Ye, YingShi Luan, Zhudi Chen, Guangyue Meng 等AAAI 2026
- AnomalyPainter: Vision-Language-Diffusion Synergy for Realistic and Diverse Unseen Industrial Anomaly SynthesisZhangyu Lai, Yilin Lu, Xinyang Li, Jianghang Lin 等AAAI 2026
相关 Paper
- Semantics-Aware Dynamic Localization and Refinement for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen 等AAAI 2023 · 被引用 35 次
- AMLRIS: Alignment-aware Masked Learning for Referring Image SegmentationTongfei Chen, Shuo Yang, Yuguang Yang, Linlin Yang 等ICLR 2026 · 被引用 1 次
- Bottom-Up and Bidirectional Alignment for Referring Expression ComprehensionLiuwu Li, Yuqi Bu, Yi CaiACM MM 2021 · 被引用 11 次
- LAVT: Language-Aware Vision Transformer for Referring Image SegmentationZhao Yang, Jiaqi Wang, Yansong Tang, Kai Chen 等CVPR 2022 · 被引用 319 次
- Two-stage Visual Cues Enhancement Network for Referring Image SegmentationYang Jiao, Zequn Jie, Weixin Luo, Jingjing Chen 等ACM MM 2021 · 被引用 24 次
