Unified Open-World Segmentation with Multi-Modal Prompts
Yang Liu, Yufei Yin, Chenchen Jing, Muzhi Zhu, Hao Chen, Yuling Xi, Bo Feng, Hao Wang, Shiyu Li, Chunhua Shen
摘要
In this work, we present COSINE, a unified open-world segmentation model that Consolidates Open-vocabulary Segmentation and IN-context sEgmentation with multimodal prompts (e.g., text and image). COSINE exploits foundation models to extract representations for an input image and corresponding multi-modal prompts, and a SegDecoder to align these representations, model their interaction, and obtain masks specified by input prompts across different granularities. In this way, COSINE overcomes architectural discrepancies, divergent learning objectives, and distinct representation learning strategies of previous pipelines for open-vocabulary segmentation and incontext segmentation. Comprehensive experiments demonstrate that COSINE has significant performance improvements in both open-vocabulary and in-context segmentation tasks. Our exploratory analyses highlight that the synergistic collaboration between using visual and textual prompts leads to significantly improved generalization over single-modality approaches. Our code is released at https://github.com/aim-uofa/COSINE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Robust Promptable Video Object SegmentationSohyun Lee, Yeho Gwon, Lukas Hoyer, Konrad Schindler 等CVPR 2026
- UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual PromptingGeonuk Kim, Minhoi Kim, Kangil Lee, Minsu Kim 等CVPR 2026
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?Tilemachos Aravanis, Vladan Stojnic, Bill Psomas, Nikos Komodakis 等CVPR 2026
- HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in VideosTingting Han, Xinsong Tao, Yufei Yin, Min Tan 等CVPR 2026
它引用的顶会 Paper41
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- SAM-CP: Marrying SAM with Composable Prompts for Versatile SegmentationPengfei Chen, Lingxi Xie, Xinyue Huo, Xuehui Yu 等ICLR 2025
- COS3D: Collaborative Open-Vocabulary 3D SegmentationRunsong Zhu, Ka-Hei Hui, Zhengzhe Liu, Qianyi Wu 等NeurIPS 2025 · 被引用 12 次
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li 等NeurIPS 2023 · 被引用 889 次
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsShiting Xiao, Rishabh Kabra, Yuhang Li, Donghyun Lee 等NeurIPS 2025 · 被引用 15 次
- Segment Anyword: Mask Prompt Inversion for Open-Set Grounded SegmentationZhihua Liu, Amrutha Saseendran, Lei Tong, Xilin He 等ICML 2025
