Refer to Any Segmentation Mask Group with Vision-Language Prompts
Shengcao Cao, Zijun Wei, Jason Kuen, Kangning Liu, Lingzhi Zhang, Jiuxiang Gu, Hyunjoon Jung, Liang-Yan Gui, Yu-Xiong Wang
Abstract
Recent image segmentation models have advanced to segment images into high-quality masks for visual entities, and yet they cannot provide comprehensive semantic understanding for complex queries based on both language and vision. This limitation reduces their effectiveness in applications that require user-friendly interactions driven by vision-language prompts. To bridge this gap, we introduce a novel task of omnimodal referring expression segmentation (ORES). In this task, a model produces a group of masks based on arbitrary prompts specified by text only or text plus reference visual entities. To address this new challenge, we propose a novel framework to "Refer to Any Segmentation Mask Group" (RAS), which augments segmentation models with complex multimodal interactions and comprehension via a mask-centric large multimodal model. For training and benchmarking ORES models, we create datasets MaskGroups-2M and MaskGroups-HQ to include diverse mask groups specified by text and reference entities. Through extensive evaluation, we demonstrate superior performance of RAS on our new ORES task, as well as classic referring expression segmentation (RES) and generalized referring expression segmentation (GRES) tasks. Project page: https://Ref2Any.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 918dd010-6684-4100-b063-1f30105b8c75Cited by top-tier papers3
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and PredictionWeitai Kang, Jason Kuen, Mengwei Ren, Zijun Wei et al.CVPR 2026 · 5 citations
- SegCompass: Exploring Interpretable Alignment with Sparse Autoencoders for Enhanced Reasoning SegmentationZhenyu Lu, Liupeng Li, Jinpeng Wang, Haoqian Kang et al.CVPR 2026 · 3 citations
- Hugging Visual Prompt and Segmentation Tokens: Consistency Learning for Fine-Grained Visual Understanding in MLLMsjing yang, Sen Yang, Boqiang Duan, Ming Dai et al.CVPR 2026
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
Related papers
- Advancing Referring Expression Segmentation Beyond Single ImageYixuan Wu, Zhao Zhang, Chi Xie, Feng Zhu et al.ICCV 2023 · 25 citations
- GSVA: Generalized Segmentation via Multimodal Large Language ModelsZhuofan Xia, Dongchen Han, Yizeng Han, Xuran Pan et al.CVPR 2024 · 42 citations
- Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationWenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo et al.CVPR 2024 · 5 citations
- Meta Compositional Referring Expression SegmentationLi Xu, Mark He Huang, Xindi Shang, Zehuan Yuan et al.CVPR 2023
- GLaMM: Pixel Grounding Large Multimodal ModelHanoona Abdul Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman M. Shaker et al.CVPR 2024 · 113 citations
