Perceptual Grouping in Contrastive Vision-Language Models
Kanchana Ranasinghe, Brandon McKinzie, Sachin Ravi, Yinfei Yang, Alexander Toshev, Jonathon Shlens
Abstract
Recent advances in zero-shot image recognition suggest that vision-language models learn generic visual representations with a high degree of semantic information that may be arbitrarily probed with natural language phrases. Understanding an image, however, is not just about understanding what content resides within an image, but importantly, where that content resides. In this work we examine how well visionlanguage models are able to understand where objects reside within an image and group together visually related parts of the imagery. We demonstrate how contemporary vision and language representation learning models based on contrastive losses and large web-based data capture limited object localization information. We propose a minimal set of modifications that results in models that uniquely learn both semantic and spatial information. We measure this performance in terms of zero-shot image recognition, unsupervised bottom-up and top-down semantic segmentations, as well as robustness analyses. We find that the resulting model achieves state-of-the-art results in terms of unsupervised segmentation, and demonstrate that the learned representations are uniquely robust to spurious correlations in datasets designed to probe the causal behavior of vision models. * Work performed as part of Apple internship. † Work performed at Apple.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b595d05d-65dd-4285-a8d7-219057c082c5Cited by top-tier papers35
- Data Filtering NetworksAlex Fang, Albin Madappally Jose, Amit Jain, Ludwig Schmidt et al.ICLR 2024 · 251 citations
- Meta CLIP 2: A Worldwide Scaling RecipeYung-Sung Chuang, Yang Li, Dong Wang, Ching-Feng Yeh et al.NeurIPS 2025 · 72 citations
- FineCLIP: Self-distilled Region-based CLIP for Better Fine-grained UnderstandingDong Jing, Xiaolong He, Yutian Luo, Nanyi Fei et al.NeurIPS 2024 · 70 citations
- Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary SegmentationLuca Barsellotti, Lorenzo Bianchi, Nicola Messina, Fabio Carrara et al.ICCV 2025 · 58 citations
- Improving fine-grained understanding in image-text pre-trainingIoana Bica, Anastasija Ilic, Matthias Bauer, Goker Erdogan et al.ICML 2024 · 53 citations
Builds on58
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- A Simple Framework for Open-Vocabulary Zero-Shot SegmentationThomas Stegmüller, Tim Lebailly, Nikola Dukic, Behzad Bozorgtabar et al.ICLR 2025
- Exploring Open-Vocabulary Semantic Segmentation from CLIP Vision Encoder Distillation OnlyJun Chen, Deyao Zhu, Guocheng Qian, Bernard Ghanem et al.ICCV 2023 · 60 citations
- GLIPv2: Unifying Localization and Vision-Language UnderstandingHaotian Zhang, Pengchuan Zhang, Xiaowei Hu, Yen-Chun Chen et al.NeurIPS 2022 · 403 citations
- [CLS] is Not Enough: Multi-Label Recognition via Patch-Level Inference and Adaptive AggregationAkang Wang, Xili Deng, Zhanxuan Hu, Yi Zhao et al.ICML 2026 · 1 citation
- Unifying Vision-Language Representation Space with Single-Tower TransformerJiho Jang, Chaerin Kong, Donghyeon Jeon, Seonhoon Kim et al.AAAI 2023 · 34 citations
