Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
Zhiwei Yang, Yucong Meng, Kexue Fu, Feilong Tang, Shuo Wang, Zhijian Song
Abstract
Weakly Supervised Semantic Segmentation (WSSS) with image-level labels aims to achieve pixel-level predictions using Class Activation Maps (CAMs). Recently, Contrastive Language-Image Pre-training (CLIP) has been introduced in WSSS. However, recent methods primarily focus on image-text alignment for CAM generation, while CLIP's potential in patch-text alignment remains unexplored. In this work, we propose ExCEL to explore CLIP's dense knowledge via a novel patch-text alignment paradigm for WSSS. Specifically, we propose Text Semantic Enrichment (TSE) and Visual Calibration (VC) modules to improve the dense alignment across both text and vision modalities. To make text embeddings semantically informative, our TSE module applies Large Language Models (LLMs) to build a dataset-wide knowledge base and enriches the text representations with an implicit attributehunting process. To mine fine-grained knowledge from visual features, our VC module first proposes Static Visual Calibration (SVC) to propagate fine-grained knowledge in a non-parametric manner. Then Learnable Visual Calibration (LVC) is further proposed to dynamically shift the frozen features towards distributions with diverse semantics. With these enhancements, ExCEL not only retains CLIP's training-free advantages but also significantly outperforms other state-of-the-art methods with much less training cost on PASCAL VOC and MS COCO. Code is available at https://github.com/zwyang6/ExCEL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bdc7af41-322b-422f-94c1-052d8526c7eeCited by top-tier papers10
- TRACE: Your Diffusion Model is Secretly an Instance Edge DetectorSanghyun Jo, Ziseok Lee, Wooyeol Lee, Jonghyun Choi et al.ICLR 2026 · 4 citations
- OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing ImageryQi Guo, Jue Wang, Yinhe Liu, Yanfei ZhongCVPR 2026 · 2 citations
- Bias-Resilient Weakly Supervised Semantic Segmentation Using Normalizing FlowsXianglin Qiu, Xiaoyang Wang, Zhen Zhang, Jimin XiaoICCV 2025 · 1 citation
- Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language ModelsZhiwei Yang, Yuanchen Wu, Nan Zhang, Yucong Meng et al.ICML 2026 · 1 citation
- SSR: Semantic and Spatial Rectification for CLIP-based Weakly Supervised SegmentationXiuli Bi, Die Xiao, Junchao Fan, Bin XiaoAAAI 2026 · 1 citation
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware PromptingYongming Rao, Wenliang Zhao, Guangyi Chen, Yansong Tang et al.CVPR 2022 · 527 citations
- Multi-class Token Transformer for Weakly Supervised Semantic SegmentationLian Xu, Wanli Ouyang, Mohammed Bennamoun, Farid Boussaïd et al.CVPR 2022 · 275 citations
- Learning Affinity from Attention: End-to-End Weakly-Supervised Semantic Segmentation with TransformersLixiang Ru, Yibing Zhan, Baosheng Yu, Bo DuCVPR 2022 · 257 citations
Related papers
- CLIP is Also an Efficient Segmenter: A Text-Driven Approach for Weakly Supervised Semantic SegmentationYuqi Lin, Minghao Chen, Wenxiao Wang, Boxi Wu et al.CVPR 2023
- Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic SegmentationXianglin Qiu, Jian Wang, Xiaolei Wang, Zhen Zhang et al.CVPR 2026
- CLIMS: Cross Language Image Matching for Weakly Supervised Semantic SegmentationJinheng Xie, Xianxu Hou, Kai Ye, Linlin ShenCVPR 2022 · 171 citations
- Toward Modality Gap: Vision Prototype Learning for Weakly-supervised Semantic Segmentation with CLIPZhongxing Xu, Feilong Tang, Zhe Chen, Yingxue Su et al.AAAI 2025 · 23 citations
- SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationHuaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He et al.ICML 2023 · 222 citations
