Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIP
Zhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su, Zhiyi Zhao, Ge Zhang, Jionglong Su, Zongyuan Ge
Abstract
The application of Contrastive Language-Image Pre-training (CLIP) in Weakly Supervised Semantic Segmentation (WSSS) research demonstrates powerful cross-modal semantic understanding capabilities. Existing methods attempt to optimize input text prompts for improved alignment of images and text, specifically by finely adjusting text prototypes to facilitate semantic matching. Nevertheless, given the modality gap between text and vision spaces, the text prototypes employed by these methods have not effectively established a close correspondence with pixel-level vision features. In this work, our theoretical analysis indicates that the inherent modality gap results in misalignment of text and region features, and that this gap cannot be sufficiently reduced by minimizing contrast loss in CLIP. To mitigate the impact of the modality gap, we propose a Vision Prototype Learning (VPL) framework, by introducing more representative vision prototypes. The core of this framework is to learn class-specific vision prototypes in vision space with the help of text prototypes, to capture high-quality localization maps. Moreover, we propose a regional semantic contrast module that contrasts regions embedding with corresponding prototypes, leading to more comprehensive and robust feature learning. Experimental results show that our proposed framework achieves state-of-the-art performance on two benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 074f939e-6047-42f8-99aa-44311ec702a6Cited by top-tier papers10
- MMRC: A Large-Scale Benchmark for Understanding Multimodal Large Language Model in Real-World ConversationHaochen Xue, Feilong Tang, Ming Hu, Yexin Liu et al.ACL 2025 · 23 citations
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware DecodingZhongxing Xu, Zhonghua Wang, Zhe Qian, Dachuan Shi et al.CVPR 2026 · 16 citations
- Class Token as Proxy: Optimal Transport-Assisted Proxy Learning for Weakly Supervised Semantic SegmentationJian Wang, Tianhong Dai, Bingfeng Zhang, Siyue Yu et al.ICCV 2025 · 2 citations
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict ResolutionBaoliang Tian, Yuxuan Si, Jilong Wang, Lingyao Li et al.AAAI 2026 · 2 citations
- SSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised SegmentationXiuli Bi, Die Xiao, Junchao Fan, Bin XiaoAAAI 2026 · 1 citation
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- Regional Semantic Contrast and Aggregation for Weakly Supervised Semantic SegmentationTianfei Zhou, Meijie Zhang, Fang Zhao, Jianwu LiCVPR 2022 · 190 citations
- CLIMS: Cross Language Image Matching for Weakly Supervised Semantic SegmentationJinheng Xie, Xianxu Hou, Kai Ye, Linlin ShenCVPR 2022 · 171 citations
- SuS-X: Training-Free Name-Only Transfer of Vision-Language ModelsVishaal Udandarao, Ankush Gupta, Samuel AlbanieICCV 2023 · 160 citations
Related papers
- Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic SegmentationXianglin Qiu, Jian Wang, Xiaolei Wang, Zhen Zhang et al.CVPR 2026
- CRIS: CLIP-Driven Referring Image SegmentationZhaoqing Wang, Yu Lu, Qiang Li, Xunqiang Tao et al.CVPR 2022 · 337 citations
- CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic SegmentationYuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu et al.CVPR 2023
- Overcoming the Pitfalls of Vision-Language Model for Image-Text RetrievalFeifei Zhang, Sijia Qu, Fan Shi, Changsheng XuACM MM 2024 · 12 citations
- Learning Multi-Modal Class-Specific Tokens for Weakly Supervised Dense Object LocalizationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd et al.CVPR 2023
