CLIP-S4: Language-Guided Self-Supervised Semantic Segmentation
Wenbin He, Suphanut Jamonnak, Liang Gou, Liu Ren
Abstract
Existing semantic segmentation approaches are often limited by costly pixel-wise annotations and predefined classes. In this work, we present CLIP-S 4 that leverages self-supervised pixel representation learning and visionlanguage models to enable various semantic segmentation tasks (e.g., unsupervised, transfer learning, languagedriven segmentation) without any human annotations and unknown class information. We first learn pixel embeddings with pixel-segment contrastive learning from different augmented views of images. To further improve the pixel embeddings and enable language-driven semantic segmentation, we design two types of consistency guided by visionlanguage models: 1) embedding consistency, aligning our pixel embeddings to the joint feature space of a pre-trained vision-language model, CLIP [34]; and 2) semantic consistency, forcing our model to make the same predictions as CLIP over a set of carefully designed target classes with both known and unknown prototypes. Thus, CLIP-S 4 enables a new task of class-free semantic segmentation where no unknown class information is needed during training. As a result, our approach shows consistent and substantial performance improvement over four popular benchmarks compared with the state-of-the-art unsupervised and languagedriven semantic segmentation methods. More importantly, our method outperforms these methods on unknown class recognition by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 51a8019e-3427-4c27-a75c-b10c878f7d72Cited by top-tier papers12
- MMA: Multi-Modal Adapter for Vision-Language ModelsLingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, Xiaohua XieCVPR 2024 · 46 citations
- CLIP as RNN: Segment Countless Visual Concepts without Training EndeavorShuyang Sun, Runjia Li, Philip Torr, Xiuye Gu et al.CVPR 2024 · 22 citations
- Learn to Rectify the Bias of CLIP for Unsupervised Semantic SegmentationJingyun Wang, Guoliang KangCVPR 2024 · 8 citations
- FATE: Feature-Adapted Parameter Tuning for Vision-Language ModelsZhengqin Xu, Zelin Peng, Xiaokang Yang, Wei ShenAAAI 2025 · 3 citations
- Language-driven Fine-grained RetrievalShijie Wang, Xin Yu, Yadan Luo, Zijian Wang et al.CVPR 2026 · 2 citations
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 1,274 citations
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 956 citations
Related papers
- Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation OnlyJun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem et al.ICCV 2023 · 60 citations
- DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary LookupZhen Qu, Xian Tao, Xinyi Gong, Shichen Qu et al.ICCV 2025 · 3 citations
- Rethinking Prior Information Generation with CLIP for Few-Shot SegmentationJin Wang, Bingfeng Zhang, Jian Pang, Honglong Chen et al.CVPR 2024 · 27 citations
- CLIPCleaner: Cleaning Noisy Labels with CLIPChen Feng, Georgios Tzimiropoulos, Ioannis PatrasACM MM 2024 · 12 citations
- S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist CaptionsSangwoo Mo, Minkyu Kim, Kyungmin Lee, Jinwoo ShinNeurIPS 2023 · 53 citations
