Prompt-Driven Referring Image Segmentation with Instance Contrasting
Chao Shang, Zichen Song, Heqian Qiu, Lanxiao Wang, Fanman Meng, Hongliang Li
摘要
Referring image segmentation (RIS) aims to segment the target referent described by natural language. Recently, large-scale pre-trained models, e.g., CLIP and SAM, have been successfully applied in many downstream tasks, but they are not well adapted to RIS task due to inter-task differences. In this paper, we propose a new prompt-driven framework named Prompt-RIS, which bridges CLIP and SAM end-to-end and transfers their rich knowledge and powerful capabilities to RIS task through prompt learning. To adapt CLIP to pixel-level task, we first propose a Cross-Modal Prompting method, which acquires more comprehensive vision-language interaction and fine-grained text-to-pixel alignment by performing bidirectional prompting. Then, the prompt-tuned CLIP generates masks, points, and text prompts for SAM to generate more accurate mask predictions. Moreover, we further propose Instance Contrastive Learning to improve the model's discriminability to different instances and robustness to diverse languages describing the same instance. Extensive experiments demonstrate that the performance of our method outperforms the state-of-the-art methods consistently in both general and open-vocabulary settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Multi-task Visual Grounding with Coarse-to-Fine Consistency ConstraintsMing Dai, Jian Li, Jiedong Zhuang, Xian Zhang 等AAAI 2025 · 被引用 23 次
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation Through Loopback SynergyMing Dai, Wenxuan Cheng, Jiang-Jiang Liu, Sen Yang 等ICCV 2025 · 被引用 5 次
- SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image SegmentationZhenjie Mao, Yuhuan Yang, Chaofan Ma, Dongsheng Jiang 等NeurIPS 2025 · 被引用 4 次
- GenMask: Adapting DiT for Segmentation via Direct Mask GenerationYuhuan Yang, Xianwei Zhuang, Yuxuan Cai, Chaofan Ma 等CVPR 2026 · 被引用 4 次
- CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ SegmentationXinlei Yu, Changmiao Wang, Hui Jin, Ahmed Elazab 等ACM MM 2025 · 被引用 3 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 被引用 1,438 次
相关 Paper
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao 等CVPR 2022 · 被引用 337 次
- Zero-shot Referring Image Segmentation with Global-Local Context FeaturesSeonghoon Yu, Paul Hongsuck Seo, Jeany SonCVPR 2023
- Curriculum Point Prompting for Weakly-Supervised Referring Image SegmentationQiyuan Dai, Sibei YangCVPR 2024
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su 等AAAI 2025 · 被引用 23 次
- CARIS: Context-Aware Referring Image SegmentationSun'ao Liu, Yiheng Zhang, Zhaofan Qiu, Hongtao Xie 等ACM MM 2023 · 被引用 34 次
