Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation
Yun Xing, Jian Kang, Aoran Xiao, Jiahao Nie, Ling Shao, Shijian Lu
Abstract
Vision-Language Pre-training has demonstrated its remarkable zero-shot recognition ability and potential to learn generalizable visual representations from language supervision. Taking a step ahead, language-supervised semantic segmentation enables spatial localization of textual inputs by learning pixel grouping solely from image-text pairs. Nevertheless, the state-of-the-art suffers from clear semantic gaps between visual and textual modality: plenty of visual concepts appeared in images are missing in their paired captions. Such semantic misalignment circulates in pre-training, leading to inferior zero-shot performance in dense predictions due to insufficient visual concepts captured in textual representations. To close such semantic gap, we propose Concept Curation (CoCu), a pipeline that leverages CLIP to compensate for the missing semantics. For each image-text pair, we establish a concept archive that maintains potential visually-matched concepts with our proposed vision-driven expansion and text-to-vision-guided ranking. Relevant concepts can thus be identified via cluster-guided sampling and fed into pre-training, thereby bridging the gap between visual and textual semantics. Extensive experiments over a broad suite of 8 segmentation benchmarks show that CoCu achieves superb zeroshot transfer performance and greatly boosts language-supervised segmentation baseline by a large margin, suggesting the value of bridging semantic gap in pretraining data. Code is available at https://github.com/xing0047/rewrite .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 469fca77-0ef1-44c0-85a7-ceca00cb169aCited by top-tier papers15
- Bridging the Gap: From Ad-hoc to Proactive Search in ConversationsChuan Meng, Francesco Tonolini, Fengran Mo, Nikolaos Aletras et al.SIGIR 2025 · 9 citations
- CorrCLIP: Reconstructing Patch Correlations in CLIP for Open-Vocabulary Semantic SegmentationDengke Zhang, Fagui Liu, Quan TangICCV 2025 · 6 citations
- Image-Text Co-Decomposition for Text-Supervised Semantic SegmentationJi-Jia Wu, Andy Chia-Hao Chang, Chieh-Yu Chuang, Chun-Pei Chen et al.CVPR 2024 · 6 citations
- CLIPeR: Hierarchically Improving Spatial Representation of CLIP for Open-Vocabulary Semantic SegmentationLin Sun, Jiale Cao, Jin Xie, Xiaoheng Jiang et al.ICCV 2025 · 5 citations
- Training-Free Class Purification for Open-Vocabulary Semantic SegmentationQi Chen, Lingxiao Yang, Yun Chen, Nailong Zhao et al.ICCV 2025 · 4 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- ReCo: Retrieve and Co-segment for Zero-shot TransferGyungin Shin, Weidi Xie, Samuel AlbanieNeurIPS 2022 · 160 citations
- Exploring Regional Clues in CLIP for Zero-Shot Semantic SegmentationYi Zhang, Meng-Hao Guo, Miao Wang, Shi-Min HuCVPR 2024 · 20 citations
- Semi-Supervised CLIP Adaptation by Enforcing Semantic and Trapezoidal ConsistencyKai Gan, Bo Ye, Min-Ling Zhang, Tong WeiICLR 2025
- SegCLIP: Patch Aggregation with Learnable Centers for Open-Vocabulary Semantic SegmentationHuaishao Luo, Junwei Bao, Youzheng Wu, Xiaodong He et al.ICML 2023 · 222 citations
- Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation OnlyJun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem et al.ICCV 2023 · 60 citations
