CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor
Shuyang Sun, Runjia Li, Philip Torr, Xiuye Gu, Siyang Li
Abstract
Existing open-vocabulary image segmentation methods re-quire a fine-tuning step on mask labels and/or image-text datasets. Mask labels are labor-intensive, which limits the number of categories in segmentation datasets. Con-sequently, the vocabulary capacity of pre-trained VLMs is severely reduced after fine-tuning. However, without fine-tuning, VLMs trained under weak image-text supervision tend to make suboptimal mask predictions. To alleviate these issues, we introduce a novel recurrent framework that progressively filters out irrelevant texts and enhances mask quality without training efforts. The recurrent unit is a two-stage segmenter built upon a frozen VLM. Thus, our model retains the VLM's broad vocabulary space and equips it with segmentation ability. Experiments show that our method outperforms not only the training-free counter-parts, but also those fine-tuned with millions of data sam-ples, and sets the new state-of-the-art records for both zero-shot semantic and referring segmentation. Concretely, we improve the current record by 28.8, 16.0, and 6.9 mloU on Pascal VOC, COCO Object, and Pascal Context.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers34
- SAM 3: Segment Anything with ConceptsNicolas Carion, Laura Gustafson, Yuan-Ting Hu, Shoubhik Debnath et al.ICLR 2026 · 1,103 citations
- Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationLuca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara et al.ICCV 2025 · 58 citations
- Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic SegmentationYunheng Li, Zhong-Yu Li, Quan-Sheng Zeng, Qibin Hou et al.ICML 2024 · 27 citations
- IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word EmphasisYuji Wang, Jingchen Ni, Yong Liu, Chun Yuan et al.AAAI 2025 · 23 citations
- Exploring the Underwater World Segmentation without Extra TrainingBingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao et al.CVPR 2026 · 18 citations
Builds on44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
Related papers
- Open-Vocabulary Semantic Segmentation with Mask-adapted CLIPFeng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li et al.CVPR 2023
- Learning Open-Vocabulary Semantic Segmentation Models From Natural Language SupervisionJilan Xu, Junlin Hou, Yuejie Zhang, Rui Feng et al.CVPR 2023
- Shatter and Gather: Learning Referring Image Segmentation with Text SupervisionDongwon Kim, Namyup Kim, Cuiling Lan, Suha KwakICCV 2023 · 29 citations
- Emergent Open-Vocabulary Semantic Segmentation from Off-the-Shelf Vision-Language ModelsJiayun Luo, Siddhesh Khandelwal, Leonid Sigal, Boyang LiCVPR 2024 · 10 citations
- MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic SegmentationKaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu et al.ICCV 2023 · 22 citations
