Mask Grounding for Referring Image Segmentation
Yong Xien Chng, Henry Zheng, Yizeng Han, Xuchong Qiu, Gao Huang
摘要
Referring Image Segmentation (RIS) is a challenging task that requires an algorithm to segment objects referred by free-form language expressions. Despite significant progress in recent years, most state-of-the-art (SOTA) methods still suffer from considerable language-image modality gap at the pixel and word level. These methods generally 1) rely on sentence-level language features for languageimage alignment and 2) lack explicit training supervision for fine-grained visual grounding. Consequently, they exhibit weak object-level correspondence between visual and language features. Without well-grounded features, prior methods struggle to understand complex expressions that require strong reasoning over relationships among multiple objects, especially when dealing with rarely used or ambiguous clauses. To tackle this challenge, we introduce a novel Mask Grounding auxiliary task that significantly improves visual grounding within language features, by explicitly teaching the model to learn fine-grained correspondence between masked textual tokens and their matching visual objects. Mask Grounding can be directly used on prior RIS methods and consistently bring improvements. Furthermore, to holistically address the modality gap, we also design a cross-modal alignment loss and an accompanying alignment module. These additions work synergistically with Mask Grounding. With all these techniques, our comprehensive approach culminates in MagNet (Maskgrounded Network), an architecture that significantly outperforms prior arts on three key benchmarks (RefCOCO, RefCOCO+ and G-Ref), demonstrating our method's effectiveness in addressing current limitations of RIS algorithms. Our code and pre-trained weights will be released. Correct Wrong diving man Correct third remote from the left (a) Fine-grained visual grounding is required to reason over complicated relationships among multiple objects.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring ModelingLinhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang 等NeurIPS 2024 · 被引用 45 次
- GSVA: Generalized Segmentation via Multimodal Large Language ModelsZhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan 等CVPR 2024 · 被引用 42 次
- Densely Connected Parameter-Efficient Tuning for Referring Image SegmentationJiaqi Huang, Zunnan Xu, Ting Liu, Yong Liu 等AAAI 2025 · 被引用 34 次
- RemoteSAM: Towards Segment Anything for Earth ObservationLiang Yao, Fan Liu, Delong Chen, Chuanyi Zhang 等ACM MM 2025 · 被引用 28 次
- Boosting Weakly Supervised Referring Image Segmentation via Progressive ComprehensionZaiquan Yang, Yuhao Liu, Jiaying Lin, Gerhard P. Hancke 等NeurIPS 2024 · 被引用 14 次
它引用的顶会 Paper41
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
相关 Paper
- Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationWenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo 等CVPR 2024 · 被引用 5 次
- LQMFormer: Language-Aware Query Mask Transformer for Referring Image SegmentationNisarg A. Shah, Vibashan VS, Vishal M. PatelCVPR 2024 · 被引用 11 次
- Zero-shot Referring Image Segmentation with Global-Local Context FeaturesSeonghoon Yu, Paul Hongsuck Seo, Jeany SonCVPR 2023
- Locate Then Segment: A Strong Pipeline for Referring Image SegmentationYa Jing, Tao Kong, Wei Wang, Liang Wang 等CVPR 2021
- Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual GroundingWenbo Chen, Zhen Xu, Ruotao Xu, Si Wu 等CVPR 2025
