Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
Yunheng Li, Zhong-Yu Li, Quan-Sheng Zeng, Qibin Hou, Ming-Ming Cheng
摘要
Pre-trained vision-language models, e.g., CLIP, have been successfully applied to zero-shot semantic segmentation. Existing CLIP-based approaches primarily utilize visual features from the last layer to align with text embeddings, while they neglect the crucial information in intermediate layers that contain rich object details. However, we find that directly aggregating the multilevel visual features weakens the zero-shot ability for novel classes. The large differences between the visual features from different layers make these features hard to align well with the text embeddings. We resolve this problem by introducing a series of independent decoders to align the multi-level visual features with the text embeddings in a cascaded way, forming a novel but simple framework named Cascade-CLIP. Our Cascade-CLIP is flexible and can be easily applied to existing zero-shot semantic segmentation methods. Experimental results show that our simple Cascade-CLIP achieves superior zero-shot performance on segmentation benchmarks, like COCO-Stuff, Pascal-VOC, and Pascal-Context. Our code is available at https://github. com/HVision-NKU/Cascade-CLIP .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual LearningLinlan Huang, Xusheng Cao, Haori Lu, Yifan Meng 等ICCV 2025 · 被引用 12 次
- Advancing Textual Prompt Learning with Anchored AttributesZheng Li, Yibing Song, Ming-Ming Cheng, Xiang Li 等ICCV 2025 · 被引用 8 次
- Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature AlignmentShi-Chen Zhang, Yunheng Li, Yu-Huan Wu, Qibin Hou 等ICCV 2025 · 被引用 8 次
- Object-Centric Refinement for Enhanced Zero-Shot SegmentationSrinivasa Rao Nandam, Sara Atito Ali, Zhenhua Feng, Josef Kittler 等ICLR 2026 · 被引用 5 次
- FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long TextBingchao Wang, Zhiwei Ning, Jianyu Ding, Xuanang Gao 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
相关 Paper
- Exploring Regional Clues in CLIP for Zero-Shot Semantic SegmentationYi Zhang, Meng-Hao Guo, Miao Wang, Shi-Min HuCVPR 2024 · 被引用 20 次
- Decoupling Zero-Shot Semantic SegmentationJian Ding, Nan Xue, Gui-Song Xia, Dengxin DaiCVPR 2022 · 被引用 255 次
- Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation OnlyJun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem 等ICCV 2023 · 被引用 60 次
- Open Vocabulary Semantic Segmentation with Patch Aligned Contrastive LearningJishnu Mukhoti, Tsung-Yu Lin, Omid Poursaeed, Rui Wang 等CVPR 2023
- Language-Driven Cross-Modal Classifier for Zero-Shot Multi-Label Image RecognitionYicheng Liu, Jie Wen, Chengliang Liu, Xiaozhao Fang 等ICML 2024 · 被引用 7 次
