CLIP-S4: Language-Guided Self-Supervised Semantic Segmentation
Wenbin He, Suphanut Jamonnak, Liang Gou, Liu Ren
摘要
Existing semantic segmentation approaches are often limited by costly pixel-wise annotations and predefined classes. In this work, we present CLIP-S 4 that leverages self-supervised pixel representation learning and visionlanguage models to enable various semantic segmentation tasks (e.g., unsupervised, transfer learning, languagedriven segmentation) without any human annotations and unknown class information. We first learn pixel embeddings with pixel-segment contrastive learning from different augmented views of images. To further improve the pixel embeddings and enable language-driven semantic segmentation, we design two types of consistency guided by visionlanguage models: 1) embedding consistency, aligning our pixel embeddings to the joint feature space of a pre-trained vision-language model, CLIP [34]; and 2) semantic consistency, forcing our model to make the same predictions as CLIP over a set of carefully designed target classes with both known and unknown prototypes. Thus, CLIP-S 4 enables a new task of class-free semantic segmentation where no unknown class information is needed during training. As a result, our approach shows consistent and substantial performance improvement over four popular benchmarks compared with the state-of-the-art unsupervised and languagedriven semantic segmentation methods. More importantly, our method outperforms these methods on unknown class recognition by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- MMA: Multi-Modal Adapter for Vision-Language ModelsLingxiao Yang, Ru-Yuan Zhang, Yanchen Wang, Xiaohua XieCVPR 2024 · 被引用 46 次
- CLIP as RNN: Segment Countless Visual Concepts without Training EndeavorShuyang Sun, Runjia Li, Philip Torr, Xiuye Gu 等CVPR 2024 · 被引用 22 次
- Learn to Rectify the Bias of CLIP for Unsupervised Semantic SegmentationJingyun Wang, Guoliang KangCVPR 2024 · 被引用 8 次
- FATE: Feature-Adapted Parameter Tuning for Vision-Language ModelsZhengqin Xu, Zelin Peng, Xiaokang Yang, Wei ShenAAAI 2025 · 被引用 3 次
- Language-driven Fine-grained RetrievalShijie Wang, Xin Yu, Yadan Luo, Zijian Wang 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- Open-vocabulary Object Detection via Vision and Language Knowledge DistillationXiuye Gu, Tsung-Yi Lin, Weicheng Kuo, Yin CuiICLR 2022 · 被引用 1,274 次
- Invariant Information Clustering for Unsupervised Image Classification and SegmentationXu Ji, Andrea Vedaldi, João F. HenriquesICCV 2019 · 被引用 956 次
相关 Paper
- Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation OnlyJun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem 等ICCV 2023 · 被引用 60 次
- DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary LookupZhen Qu, Xian Tao, Xinyi Gong, Shichen Qu 等ICCV 2025 · 被引用 3 次
- Rethinking Prior Information Generation with CLIP for Few-Shot SegmentationJin Wang, Bingfeng Zhang, Jian Pang, Honglong Chen 等CVPR 2024 · 被引用 27 次
- CLIPCleaner: Cleaning Noisy Labels with CLIPChen Feng, Georgios Tzimiropoulos, Ioannis PatrasACM MM 2024 · 被引用 12 次
- S-CLIP: Semi-supervised Vision-Language Learning using Few Specialist CaptionsSangwoo Mo, Minkyu Kim, Kyungmin Lee, Jinwoo ShinNeurIPS 2023 · 被引用 53 次
