Exploring Regional Clues in CLIP for Zero-Shot Semantic Segmentation
Yi Zhang, Meng-Hao Guo, Miao Wang, Shi-Min Hu
Abstract
CLIP has demonstrated marked progress in visual recognition due to its powerful pre-training on large-scale image-text pairs. However, it still remains a critical challenge: how to transfer image-level knowledge into pixel-level understanding tasks such as semantic segmentation. In this paper, to solve the mentioned challenge, we analyze the gap between the capability of the CLIP model and the requirement of the zero-shot semantic segmentation task. Based on our analysis and observations, we propose a novel method for zero-shot semantic segmentation, dubbed CLIP-RC (CLIP with Regional Clues), bringing two main insights. On the one hand, a region-level bridge is necessary to provide fine-grained semantics. On the other hand, over-fitting should be mitigated during the training stage. Benefiting from the above discoveries, CLIP-RC achieves state-of-the-art performance on various zero-shot semantic segmentation benchmarks, including PASCAL VOC, PASCAL Context, and COCO-Stuff 164K. Code will be available at https://github.com/Jittor/JSeg.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 091f3e82-27fe-4c21-a52b-e5ac6d919b34Cited by top-tier papers9
- Exploring Semantic Consistency and Style Diversity for Domain Generalized Semantic SegmentationHongwei Niu, Linhuang Xie, Jianghang Lin, Shengchuan ZhangAAAI 2025 · 16 citations
- Object-Centric Refinement for Enhanced Zero-Shot SegmentationSrinivasa Rao Nandam, Sara Atito Ali, Zhenhua Feng, Josef Kittler et al.ICLR 2026 · 5 citations
- Probabilistic Prototype Calibration of Vision-Language Models for Generalized Few-Shot Semantic SegmentationJie Liu, Jiayi Shen, Pan Zhou, Jan-Jakob Sonke et al.ICCV 2025 · 4 citations
- FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long TextBingchao Wang, Zhiwei Ning, Jianyu Ding, Xuanang Gao et al.ICCV 2025 · 3 citations
- Beyond Text: Visual Description Assembly by Probabilistic Model for CLIP-based Weakly Supervised Semantic SegmentationXianglin Qiu, Jian Wang, Xiaolei Wang, Zhen Zhang et al.CVPR 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Conditional Prompt Learning for Vision-Language ModelsKaiyang Zhou, Jingkang Yang, Chen Change Loy, Ziwei LiuCVPR 2022 · 1,438 citations
Related papers
- ZegCLIP: Towards Adapting CLIP for Zero-shot Semantic SegmentationZiqin Zhou, Yinjie Lei, Bowen Zhang, Lingqiao Liu et al.CVPR 2023
- Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic SegmentationYunheng Li, Zhong-Yu Li, Quan-Sheng Zeng, Qibin Hou et al.ICML 2024 · 27 citations
- Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive LearningJishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang et al.CVPR 2023
- Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic SegmentationYun Xing, Jian Kang, Aoran Xiao, Jiahao Nie et al.NeurIPS 2023 · 29 citations
- RegionCLIP: Region-based Language-Image PretrainingYiwu Zhong, Jianwei Yang, Pengchuan Zhang, Chunyuan Li et al.CVPR 2022 · 481 citations
