Relationship Prompt Learning is Enough for Open-Vocabulary Semantic Segmentation
Jiahao Li, Yang Lu, Yuan Xie, Yanyun Qu
摘要
Open-vocabulary semantic segmentation (OVSS) aims to segment unseen classes without corresponding labels. Existing Vision-Language Model (VLM)- based methods leverage VLM’s rich knowledge to enhance additional explicit segmentation-specific networks, yielding competitive results, but at the cost of extensive training cost. To reduce the cost, we attempt to enable VLM to directly produce the segmentation results without any segmentation-specific networks. Prompt learning offers a direct and parameter-efficient approach, yet it falls short in guiding VLM for pixel-level visual classification. Therefore, we propose the R elationship P rompt M odule ( RPM ), which generates the relationship prompt that directs VLM to extract pixel-level semantic embeddings suitable for OVSS. More-over, RPM integrates with VLM to construct the R elationship P rompt N etwork ( RPN ), achieving OVSS without any segmentation-specific networks. RPN attains state-of-the-art performance with merely about 3M trainable parameters (2% of total parameters).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Target Refocusing via Attention Redistribution for Open-Vocabulary Semantic Segmentation: An Explainability PerspectiveJiahao Li, Yang Lu, Yachao Zhang, Yong Xie 等AAAI 2026 · 被引用 3 次
- Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual VariationsYiwen Liang, Hui Chen, Yizhe Xiong, Zihan Zhou 等ACM MM 2025 · 被引用 1 次
- From Words to Pixels: A Comprehensive Survey on Large Language Models in Visual SegmentationYizhou Wang, Mang Tik Chiu, Lingzhi Zhang, Xuan Shen 等ACL 2026
- Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic SegmentationJiahao Li, Yang Lu, Yachao Zhang, Fangyong Wang 等CVPR 2026
- Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic SegmentationJiahao Li, Yang Lu, Yachao Zhang, Fangyong Wang 等ACM MM 2025
它引用的顶会 Paper42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Do Vision Transformers See Like Convolutional Neural Networks?Maithra Raghu, Thomas Unterthiner, Simon Kornblith, Chiyuan Zhang 等NeurIPS 2021 · 被引用 1,553 次
相关 Paper
- Emergent Open-Vocabulary Semantic Segmentation from Off-the-Shelf Vision-Language ModelsJiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang LiCVPR 2024 · 被引用 10 次
- OpenWorldSAM: Extending SAM2 for Universal Image Segmentation with Language PromptsShiting Xiao, Rishabh Kabra, Yuhang Li, Donghyun Lee 等NeurIPS 2025 · 被引用 15 次
- LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic SegmentationVladan Stojnic, Yannis Kalantidis, Jirí Matas, Giorgos ToliasCVPR 2025
- Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic SegmentationChanyoung Kim, Dayun Ju, Woojung Han, Ming-Hsuan Yang 等CVPR 2025
- Prompt Pre-Training with Twenty-Thousand Classes for Open-Vocabulary Visual RecognitionShuhuai Ren, Aston Zhang, Yi Zhu, Shuai Zhang 等NeurIPS 2023 · 被引用 46 次
