Relationship Prompt Learning is Enough for Open-Vocabulary Semantic Segmentation
Jiahao Li, Yang Lu, Yuan Xie, Yanyun Qu
Abstract
Open-vocabulary semantic segmentation (OVSS) aims to segment unseen classes without corresponding labels. Existing Vision-Language Model (VLM)- based methods leverage VLM’s rich knowledge to enhance additional explicit segmentation-specific networks, yielding competitive results, but at the cost of extensive training cost. To reduce the cost, we attempt to enable VLM to directly produce the segmentation results without any segmentation-specific networks. Prompt learning offers a direct and parameter-efficient approach, yet it falls short in guiding VLM for pixel-level visual classification. Therefore, we propose the R elationship P rompt M odule ( RPM ), which generates the relationship prompt that directs VLM to extract pixel-level semantic embeddings suitable for OVSS. More-over, RPM integrates with VLM to construct the R elationship P rompt N etwork ( RPN ), achieving OVSS without any segmentation-specific networks. RPN attains state-of-the-art performance with merely about 3M trainable parameters (2% of total parameters).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78d40421-87f9-4d3c-ba20-6d94a6c9f62aCited by top-tier papers5
- Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability PerspectiveJiahao Li, Yang Lu, Yachao Zhang, Yong Xie et al.AAAI 2026 · 3 citations
- Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual VariationsYiwen Liang, Hui Chen, Yizhe Xiong, Zihan Zhou et al.ACM MM 2025 · 1 citation
- From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual SegmentationYizhou Wang, Mang Tik Chiu, Lingzhi Zhang, Xuan Shen et al.ACL 2026
- Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic SegmentationJiahao Li, Yang Lu, Yachao Zhang, Fangyong Wang et al.CVPR 2026
- Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic SegmentationJiahao Li, Yang Lu, Yachao Zhang, Fangyong Wang et al.ACM MM 2025
Builds on42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang et al.NeurIPS 2021 · 1,553 citations
Related papers
- Emergent Open-Vocabulary Semantic Segmentation from Off-the-Shelf Vision-Language ModelsJiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang LiCVPR 2024 · 10 citations
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsShiting Xiao, Rishabh Kabra, Yuhang Li, Donghyun Lee et al.NeurIPS 2025 · 15 citations
- LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic SegmentationVladan Stojnic, Yannis Kalantidis, Jirí Matas, Giorgos ToliasCVPR 2025
- Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic SegmentationChanyoung Kim, Dayun Ju, Woojung Han, Ming-Hsuan Yang et al.CVPR 2025
- Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual RecognitionShuhuai Ren, Aston Zhang, Yi Zhu, Shuai Zhang et al.NeurIPS 2023 · 46 citations
