LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic Segmentation
Vladan Stojnic, Yannis Kalantidis, Jirí Matas, Giorgos Tolias
Abstract
We propose a training-free method for open-vocabulary semantic segmentation using Vision-and-Language Models (VLMs). Our approach enhances the initial per-patch predictions of VLMs through label propagation, which jointly optimizes predictions by incorporating patch-to-patch relationships. Since VLMs are primarily optimized for crossmodal alignment and not for intra-modal similarity, we use a Vision Model (VM) that is observed to better capture these relationships. We address resolution limitations inherent to patch-based encoders by applying label propagation at the pixel level as a refinement step, significantly improving segmentation accuracy near class boundaries. Our method, called LPOSS+, performs inference over the entire image, avoiding window-based processing and thereby capturing contextual interactions across the full image. LPOSS+ achieves state-of-the-art performance among training-free methods, across a diverse set of datasets. Code: https: //github.com/vladan-stojnic/LPOSS
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16bc51e5-85f9-4045-8a0d-7d878f810187Cited by top-tier papers10
- PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic SegmentationGensheng Pei, Xiruo Jiang, Xinhao Cai, Tao Chen et al.CVPR 2026 · 3 citations
- OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance InformationXuehui Wang, Chongjie Si, Xue Yang, Yuzhi Zhao et al.NeurIPS 2025 · 3 citations
- Semi-Supervised Semantic Segmentation via Derivative Label PropagationYuanbin Fu, Xiaojie GuoAAAI 2026 · 1 citation
- Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation ModelsShuwen Yu, Zhanxuan Hu, Yi Zhao, Yonghang Tai et al.ICML 2026 · 1 citation
- Noise Self-Correction via Relation Propagation for Robust Cross-Modal RetrievalRuoxuan Li, Xiangyu Wu, Yang YangACM MM 2025
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- Emergent Open-Vocabulary Semantic Segmentation from Off-the-Shelf Vision-Language ModelsJiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang LiCVPR 2024 · 10 citations
- Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic SegmentationChanyoung Kim, Dayun Ju, Woojung Han, Ming-Hsuan Yang et al.CVPR 2025
- Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region MatchingXuanpu Zhao, Dianmo Sheng, Zhentao Tan, Zhiwei Zhao et al.AAAI 2025 · 2 citations
- Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic SegmentationYajie Liu, Guodong Wang, Jinjin Zhang, Qingjie Liu et al.AAAI 2025 · 3 citations
- Relationship Prompt Learning is Enough for Open-Vocabulary Semantic SegmentationJiahao Li, Yang Lu, Yuan Xie, Yanyun QuNeurIPS 2024 · 12 citations
