CLIMS: Cross Language Image Matching for Weakly Supervised Semantic Segmentation
Jinheng Xie, Xianxu Hou, Kai Ye, Linlin Shen
Abstract
It has been widely known that CAM (Class Activation Map) usually only activates discriminative object regions and falsely includes lots of object-related backgrounds. As only a fixed set of image-level object labels are available to the WSSS (weakly supervised semantic segmentation) model, it could be very difficult to suppress those diverse background regions consisting of open set objects. In this paper, we propose a novel Cross Language Image Matching (CLIMS) framework, based on the recently introduced Contrastive Language-Image Pre-training (CLIP) model, for WSSS. The core idea of our framework is to introduce natural language supervision to activate more complete object regions and suppress closely-related open background regions. In particular, we design object, background region and text label matching losses to guide the model to excite more reasonable object regions for CAM of each category. In addition, we design a co-occurring background suppression loss to prevent the model from activating closely-related background regions, with a predefined set of class-related background text descriptions. These designs enable the proposed CLIMS to generate a more complete and compact activation map for the target objects. Extensive experiments on PASCAL VOC2012 dataset show that our CLIMS significantly outperforms the previous state-of-the-art methods. Code will be available at https://github.com/CVI-SZU/CLIMS.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16a8904e-2f26-400d-aa6f-ad5052c3c559Cited by top-tier papers40
- SFC: Shared Feature Calibration in Weakly Supervised Semantic SegmentationXinqiao Zhao, Feilong Tang, Xiaoyang Wang, Jimin XiaoAAAI 2024 · 66 citations
- FPR: False Positive Rectification for Weakly Supervised Semantic SegmentationLiyi Chen, Chenyang Lei, Ruihuang Li, Shuai Li et al.ICCV 2023 · 65 citations
- Referring Image Segmentation Using Text SupervisionFang Liu, Yuhao Liu, Yuqiu Kong, Ke Xu et al.ICCV 2023 · 52 citations
- Hunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic SegmentationFeilong Tang, Zhongxing Xu, Zhaojun Qu, Wei Feng et al.CVPR 2024 · 41 citations
- TagCLIP: A Local-to-Global Framework to Enhance Open-Vocabulary Multi-Label Classification of CLIP without TrainingYuqi Lin, Minghao Chen, Kaipeng Zhang, Hengjia Li et al.AAAI 2024 · 39 citations
Builds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Integral Object Mining via Online Attention AccumulationPeng-Tao Jiang, Qibin Hou, Yang Cao, Ming-Ming Cheng et al.ICCV 2019 · 246 citations
- Leveraging Auxiliary Tasks with Affinity Learning for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd et al.ICCV 2021 · 152 citations
- Unlocking the Potential of Ordinary Classifier: Class-specific Adversarial Erasing Framework for Weakly Supervised Semantic SegmentationHyeokjun Kweon, Sung-Hoon Yoon, Hyeonseong Kim, Daehee Park et al.ICCV 2021 · 151 citations
- Discriminative Region Suppression for Weakly-Supervised Semantic SegmentationBeomyoung Kim, Sangeun Han, Junmo KimAAAI 2021 · 137 citations
Related papers
- QA-CLIMS: Question-Answer Cross Language Image Matching for Weakly Supervised Semantic SegmentationSonghe Deng, Wei Zhuo, Jinheng Xie, Linlin ShenACM MM 2023 · 12 citations
- CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic SegmentationYuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu et al.CVPR 2023
- Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic SegmentationXianglin Qiu, Jian Wang, Xiaolei Wang, Zhen Zhang et al.CVPR 2026
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su et al.AAAI 2025 · 23 citations
- SSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised SegmentationXiuli Bi, Die Xiao, Junchao Fan, Bin XiaoAAAI 2026 · 1 citation
