LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic Segmentation
Vladan Stojnic, Yannis Kalantidis, Jirí Matas, Giorgos Tolias
摘要
We propose a training-free method for open-vocabulary semantic segmentation using Vision-and-Language Models (VLMs). Our approach enhances the initial per-patch predictions of VLMs through label propagation, which jointly optimizes predictions by incorporating patch-to-patch relationships. Since VLMs are primarily optimized for crossmodal alignment and not for intra-modal similarity, we use a Vision Model (VM) that is observed to better capture these relationships. We address resolution limitations inherent to patch-based encoders by applying label propagation at the pixel level as a refinement step, significantly improving segmentation accuracy near class boundaries. Our method, called LPOSS+, performs inference over the entire image, avoiding window-based processing and thereby capturing contextual interactions across the full image. LPOSS+ achieves state-of-the-art performance among training-free methods, across a diverse set of datasets. Code: https: //github.com/vladan-stojnic/LPOSS
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- PEARL: Geometry Aligns Semantics for Training-Free Open-Vocabulary Semantic SegmentationGensheng Pei, Xiruo Jiang, Xinhao Cai, Tao Chen 等CVPR 2026 · 被引用 3 次
- OPMapper: Enhancing Open-Vocabulary Semantic Segmentation with Multi-Guidance InformationXuehui Wang, Chongjie Si, Xue Yang, Yuzhi Zhao 等NeurIPS 2025 · 被引用 3 次
- Semi-Supervised Semantic Segmentation via Derivative Label PropagationYuanbin Fu, Xiaojie GuoAAAI 2026 · 被引用 1 次
- Geometry-Preserving Unsupervised Alignment for Heterogeneous Foundation ModelsShuwen Yu, Zhanxuan Hu, Yi Zhao, Yonghang Tai 等ICML 2026 · 被引用 1 次
- Noise Self-Correction via Relation Propagation for Robust Cross-Modal RetrievalRuoxuan Li, Xiangyu Wu, Yang YangACM MM 2025
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
相关 Paper
- Emergent Open-Vocabulary Semantic Segmentation from Off-the-Shelf Vision-Language ModelsJiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang LiCVPR 2024 · 被引用 10 次
- Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic SegmentationChanyoung Kim, Dayun Ju, Woojung Han, Ming-Hsuan Yang 等CVPR 2025
- Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region MatchingXuanpu Zhao, Dianmo Sheng, Zhentao Tan, Zhiwei Zhao 等AAAI 2025 · 被引用 2 次
- Unveiling the Knowledge of CLIP for Training-Free Open-Vocabulary Semantic SegmentationYajie Liu, Guodong Wang, Jinjin Zhang, Qingjie Liu 等AAAI 2025 · 被引用 3 次
- Relationship Prompt Learning is Enough for Open-Vocabulary Semantic SegmentationJiahao Li, Yang Lu, Yuan Xie, Yanyun QuNeurIPS 2024 · 被引用 12 次
