USE: Universal Segment Embeddings for Open-Vocabulary Image Segmentation
Xiaoqi Wang, Wenbin He, Xiwei Xuan, Clint Sebastian, Jorge Piazentin Ono, Xin Li, Sima Behpour, Thang Doan, Liang Gou, Han-Wei Shen, Liu Ren
Abstract
The open-vocabulary image segmentation task involves partitioning images into semantically meaningful segments and classifying them with flexible text-defined categories. The recent vision-based foundation models such as the Segment Anything Model (SAM) have shown superior performance in generating class-agnostic image segments. The main challenge in open-vocabulary image segmentation now lies in accurately classifying these segments into text-defined categories. In this paper, we introduce the Universal Segment Embedding (USE) framework to address this challenge. This framework is comprised of two key components: 1) a data pipeline designed to efficiently curate a large amount of segment-text pairs at various granularities, and 2) a universal segment embedding model that enables precise segment classification into a vast range of text-defined categories. The USE model can not only help open-vocabulary image segmentation but also facilitate other downstream tasks (e.g., querying and ranking). Through comprehensive experimental studies on semantic segmentation and part segmentation benchmarks, we demonstrate that the USE framework outperforms state-of-the-art open-vocabulary segmentation methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c7a5406-a567-482e-9dd2-365fb43de43bCited by top-tier papers7
- CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic SegmentationDengke Zhang, Fagui Liu, Quan TangICCV 2025 · 6 citations
- ReME: A Data-Centric Framework for Training-Free Open-Vocabulary SegmentationXiwei Xuan, Ziquan Deng, Kwan-Liu MaICCV 2025 · 3 citations
- ProSAM: Enhancing the Robustness of Sam-Based Visual Reference Segmentation with Probabilistic PromptsXiaoqi Wang, Clint Sebastian, Wenbin He, Liu RenICCV 2025 · 1 citation
- AcZeroTS: Active Learning for Zero-Shot Tissue Segmentation in Pathology ImagesJiao Tang, Junjie Zhou, Bo Qian, Peng Wan et al.ICCV 2025 · 1 citation
- Simple-ViLMedSAM: Simple Text Prompts Meet Vision-Language Models for Medical Image SegmentationChengcan Qian, Dong Nie, Geng Chen, Daoqiang Zhang et al.CVPR 2026
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
Related papers
- SAM-CP: Marrying SAM with Composable Prompts for Versatile SegmentationPengfei Chen, Lingxi Xie, Xinyue Huo, Xuehui Yu et al.ICLR 2025
- Exploring Simple Open-Vocabulary Semantic SegmentationZihang LaiCVPR 2025
- Auto-Vocabulary Semantic SegmentationOsman Ülger, Maksymilian Kulicki, Yuki Asano, Martin R. OswaldICCV 2025 · 3 citations
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsShiting Xiao, Rishabh Kabra, Yuhang Li, Donghyun Lee et al.NeurIPS 2025 · 15 citations
- Open-Vocabulary 3D Semantic Segmentation with Foundation ModelsLi Jiang, Shaoshuai Shi, Bernt SchieleCVPR 2024
