Multi-Label Prototype Visual Spatial Search for Weakly Supervised Semantic Segmentation
Songsong Duan, Xi Yang, Nannan Wang
Abstract
Existing Weakly Supervised Semantic Segmentation (WSSS) relies on the CNN-based Class Activation Map (CAM) and Transformer-based self-attention map to generate classspecific masks for semantic segmentation. However, CAM and self-attention maps usually cause incomplete segmentation due to classification bias issue. To address this issue, we propose a Multi-Label Prototype Visual Spatial Search (MuP-VSS) method with a spatial query mechanism. Specifically, MuP-VSS consists of two key components: multilabel prototype representation and multi-label prototype optimization. The former designs a global embedding to learn the global tokens from the images, and then proposes a Prototype Embedding Module (PEM) to interact with patch tokens to understand the local semantic information. The latter utilizes the exclusivity and consistency principles of the multi-label prototypes to design three prototype losses to optimize them, which contain cross-class prototype (CCP) contrastive loss, cross-image prototype (CIP) contrastive loss, and patch-to-prototype (P2P) consistency loss. CCP loss models exclusivity of multi-label prototypes learned from a single image to enhance the discriminative properties of each class better. CCP loss learns the consistency of the same class-specific prototypes extracted from multiple images to enhance the semantic consistency. P2P loss is proposed to control the semantic response of the prototype to the image patches. Experimental results on Pascal VOC 2012 and MS COCO show that MuP-VSS significantly outperforms recent methods and achieves state-of-the-art performance. * Corresponding author ๐๐ โ โ ๐ซ๐ซร๐ฏ๐ฏร๐พ๐พ Class Weight ๐ฆ๐ฆ โ โ ๐ซ๐ซร๐ช๐ช (a) CNN-based Methods Similarity Scores (c) Our Query-based MuP-VSS ๐๐ โ โ ๐ฏ๐ฏ๐พ๐พร๐ซ๐ซ ๐๐ โ โ ๐ช๐ชร๐ซ๐ซ Multi-label Tokens ๐๐ ๐๐ ๐๐ Search Encoder Encoder Channel Aggregation ๐๐ โ โ ๐ซ๐ซร๐ฏ๐ฏร๐พ๐พ Encoder Conv Class-aware Feature (b) Transformer-based Methods Refine Patch Attention Maps
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e45ea24d-4d5c-4ed0-84e9-52861ec8a5d3Cited by top-tier papers3
- Controllable-Lpmoe: Adapting to Challenging Object Segmentation Via Dynamic Local Priors From Mixture-Of-ExpertsYanguang Sun, Jiawei Lian, Jian Yang, Lei LuoICCV 2025 ยท 4 citations
- Leveraging Class Distributions in CLIP for Weakly Supervised Semantic SegmentationZiqian Yang, Xinqiao Zhao, Xiaolei Wang, Quan Zhang et al.CVPR 2026
- Frequency-Aware Affinity for Weakly Supervised Semantic SegmentationZiqian Yang, Xianglin Qiu, Xinqiao Zhao, Xiaolei Wang et al.CVPR 2026
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 ยท 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 ยท 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 ยท 13,211 citations
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaรฏd et al.CVPR 2022 ยท 275 citations
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object LocalizationWei Gao, Fang Wan, Xingjia Pan, Zhiliang Peng et al.ICCV 2021 ยท 260 citations
Related papers
- Hunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic SegmentationFeilong Tang, Zhongxing Xu, Zhaojun Qu, Wei Feng et al.CVPR 2024 ยท 41 citations
- Self-supervised Image-specific Prototype Exploration for Weakly Supervised Semantic SegmentationQi Chen, Lingxiao Yang, Jianhuang Lai, Xiaohua XieCVPR 2022 ยท 182 citations
- Class Token as Proxy: Optimal Transport-Assisted Proxy Learning for Weakly Supervised Semantic SegmentationJian Wang, Tianhong Dai, Bingfeng Zhang, Siyue Yu et al.ICCV 2025 ยท 2 citations
- Token Contrast for Weakly-Supervised Semantic SegmentationLixiang Ru, Heliang Zheng, Yibing Zhan, Bo DuCVPR 2023
- POT: Prototypical Optimal Transport for Weakly Supervised Semantic SegmentationJian Wang, Tianhong Dai, Bingfeng Zhang, Siyue Yu et al.CVPR 2025
