Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic Segmentation
Xianglin Qiu, Jian Wang, Xiaolei Wang, Zhen Zhang, Jimin Xiao
Abstract
Contrastive Language-Image Pre-training (CLIP) offers a new paradigm for Weakly Supervised Semantic Segmentation (WSSS) by generating Class Activation Maps (CAMs) from text-image alignment. Existing methods primarily rely on hand-crafted templates or general attribute descriptions generated by a large language model to construct text prototypes for querying visual features. However, these strategies faces two major limitations: the inherent modality gap in CLIP prevents text prototypes achieving tight alignment with visual features; and their static text prototypes cannot adaptively respond to target instances that exhibit diverse visual attributes. To address these challenges, our key insight is to directly construct instance-specific visual description prototype as query, thereby bypassing the suboptimal static text description optimization. To this end, we propose the Visual Description Assembly (VDA) framework. It employs a probabilistic model to map complex CLIP visual features into a structured latent space. This latent space allows us to explicitly disentangle and aggregate varied visual attributes, and then dynamically assemble them into instance-specific visual prototypes. Furthermore, to enhance the robustness of this prototype, we adaptively incorporate the semantically stable text prototype into it as the final query for generating superior CAMs. Experimental results show our method outperforms existing baselines, achieving state-of-the-art performance on WSSS benchmarks. Code is available at here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 386f8e68-6b24-45a0-bd9d-3771ca872cd6Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd et al.CVPR 2022 · 275 citations
- Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with TransformersLixiang Ru, Yibing Zhan, Baosheng Yu, Bo DuCVPR 2022 · 257 citations
- Reliability Does Matter: An End-to-End Weakly Supervised Semantic Segmentation ApproachBingfeng Zhang, Jimin Xiao, Yunchao Wei, Mingjie Sun et al.AAAI 2020 · 227 citations
- Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic SegmentationTianfei Zhou, Meijie Zhang, Fang Zhao, Jianwu LiCVPR 2022 · 190 citations
Related papers
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su et al.AAAI 2025 · 23 citations
- CLIMS: Cross Language Image Matching for Weakly Supervised Semantic SegmentationJinheng Xie, Xianxu Hou, Kai Ye, Linlin ShenCVPR 2022 · 171 citations
- Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic SegmentationZhiwei Yang, Yucong Meng, Kexue Fu, Feilong Tang et al.CVPR 2025
- CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic SegmentationYuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu et al.CVPR 2023
- A Simple Framework for Text-Supervised Semantic SegmentationMuyang Yi, Quan Cui, Hao Wu, Cheng Yang et al.CVPR 2023
