Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region Matching
Xuanpu Zhao, Dianmo Sheng, Zhentao Tan, Zhiwei Zhao, Tao Gong, Qi Chu, Bin Liu, Nenghai Yu
Abstract
Open-vocabulary semantic segmentation (OVSS) aims to segment images of arbitrary categories specified by class labels. While previous approaches relied on extensive imagetext pairs or dense semantic annotations, recent training-free methods attempted to overcome these limitations by constructing semantic prototypes in the construction stage and image-to-image matching (i.e., prototype matching) during testing. However, these methods often struggle to effectively capture the visual characteristics of categories and fail to utilize local features during prototype matching. To deal with these problems, we propose a novel training-free framework for OVSS that constructs diverse prototypes and performs fine-grained sub-region matching. Specifically, our method leverages Large Language Models (LLMs) to guide support image generation by descriptions of different attributes of categories and employs coarse-fine clustering to obtain diverse and robust part-level prototypes in the construction stage. During testing, we propose a sub-region matching method, which assigns part-level prototypes to sub-regions utilizing optimal transport, to fully utilize local image features among part-level prototypes. Extensive experiments demonstrate the effectiveness of our method and show that our method achieves state-of-the-art performance, outperforming previous methods across five datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMsXuanpu Zhao, Zhentao Tan, Dianmo Sheng, Tianxiang Chen et al.CVPR 2026 · 1 citation
- CDICS: Delving Into Fine-Grained Attribute for In-Context Segmentation via Compositional Prompts and Phased DecouplingZhiyu Li, Dianmo Sheng, Qi Chu, Shilong Chen et al.CVPR 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Prompt-to-Prompt Image Editing with Cross-Attention ControlAmir Hertz, Ron Mokady, Jay Tenenbaum, Kfir Aberman et al.ICLR 2023 · 361 citations
Related papers
- Emergent Open-Vocabulary Semantic Segmentation from Off-the-Shelf Vision-Language ModelsJiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang LiCVPR 2024 · 10 citations
- Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic SegmentationChanyoung Kim, Dayun Ju, Woojung Han, Ming-Hsuan Yang et al.CVPR 2025
- LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic SegmentationVladan Stojnic, Yannis Kalantidis, Jirí Matas, Giorgos ToliasCVPR 2025
- Direct Segmentation without Logits Optimization for Training-Free Open-Vocabulary Semantic SegmentationJiahao Li, Yang Lu, Yachao Zhang, Fangyong Wang et al.CVPR 2026
- Mask-Free OVIS: Open-Vocabulary Instance Segmentation without Manual Mask AnnotationsVibashan VS, Ning Yu, Chen Xing, Can Qin et al.CVPR 2023
