Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation Only
Jun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem, Zhicheng Yan, Chenchen Zhu, Fanyi Xiao, Sean Chang Culatana, Mohamed Elhoseiny
Abstract
Semantic segmentation is a crucial task in computer vision that involves segmenting images into semantically meaningful regions at the pixel level. However, existing approaches often rely on expensive human annotations as supervision for model training, limiting their scalability to large, unlabeled datasets. To address this challenge, we present ZeroSeg, a novel method that leverages the existing pretrained vision-language (VL) model (e.g. CLIP vision encoder [39] ) to train open-vocabulary zero-shot semantic segmentation models. Although acquired extensive knowledge of visual concepts, it is non-trivial to exploit knowledge from these VL models to the task of semantic segmentation, as they are usually trained at an image level. ZeroSeg overcomes this by distilling the visual concepts learned by VL models into a set of segment tokens, each summarizing a localized region of the target image. We evaluate ZeroSeg on multiple popular segmentation benchmarks, including PASCAL VOC 2012, PASCAL Context, and COCO, in a zero-shot manner Our approach achieves state-of-the-art performance when compared to other zero-shot segmentation methods under the same training data, while also performing competitively compared to strongly supervised methods. Finally, we also demonstrated the effectiveness of ZeroSeg on open-vocabulary segmentation, through both human studies and qualitative visualizations. The code is publicly available at https: //github.com/facebookresearch/ZeroSeg
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e0cd1c4d-8d0f-4f21-b40a-77ce8737422bCited by top-tier papers17
- CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense PredictionSize Wu, Wenwei Zhang, Lumin Xu, Sheng Jin et al.ICLR 2024 · 129 citations
- CLIP-CID: Efficient CLIP Distillation via Cluster-Instance DiscriminationKaicheng Yang, Tiancheng Gu, Xiang An, Haiqiang Jiang et al.AAAI 2025 · 26 citations
- Image-to-Image Matching via Foundation Models: A New Perspective for Open-Vocabulary Semantic SegmentationYuan Wang, Rui Sun, Naisong Luo, Yuwen Pan et al.CVPR 2024 · 13 citations
- Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic SegmentationMehrdad Noori, David Osowiechi, Gustavo Adolfo Vargas Hakim, Ali Bahri et al.NeurIPS 2025 · 13 citations
- Emergent Open-Vocabulary Semantic Segmentation from Off-the-Shelf Vision-Language ModelsJiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang LiCVPR 2024 · 10 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Decoupling Zero-Shot Semantic SegmentationJian Ding, Nan Xue, Gui-Song Xia, Dengxin DaiCVPR 2022 · 255 citations
- Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic SegmentationYunheng Li, Zhong-Yu Li, Quan-Sheng Zeng, Qibin Hou et al.ICML 2024 · 27 citations
- Exploring Regional Clues in CLIP for Zero-Shot Semantic SegmentationYi Zhang, Meng-Hao Guo, Miao Wang, Shi-Min HuCVPR 2024 · 20 citations
- SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationHuaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He et al.ICML 2023 · 222 citations
- Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic SegmentationYun Xing, Jian Kang, Aoran Xiao, Jiahao Nie et al.NeurIPS 2023 · 29 citations
