Linear Spaces of Meanings: Compositional Structures in Vision-Language Models
Matthew Trager, Pramuditha Perera, Luca Zancato, Alessandro Achille, Parminder Bhatia, Stefano Soatto
Abstract
We investigate compositional structures in data embeddings from pre-trained vision-language models (VLMs). Traditionally, compositionality has been associated with algebraic operations on embeddings of words from a preexisting vocabulary. In contrast, we seek to approximate representations from an encoder as combinations of a smaller set of vectors in the embedding space. These vectors can be seen as "ideal words" for generating concepts directly within embedding space of the model. We first present a framework for understanding compositional structures from a geometric perspective. We then explain what these compositional structures entail probabilistically in the case of VLM embeddings, providing intuitions for why they arise in practice. Finally, we empirically explore these structures in CLIP’s embeddings and we evaluate their usefulness for solving different vision-language tasks such as classification, debiasing, and retrieval. Our results show that simple linear algebraic operations on embedding vectors can be used as compositional and interpretable methods for regulating the behavior of VLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef6c06e9-a018-4711-9062-67af8ca74b9aCited by top-tier papers29
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- On the Origins of Linear Representations in Large Language ModelsYibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam et al.ICML 2024 · 68 citations
- From Causal to Concept-Based Representation LearningGoutham Rajendran, Simon Buchholz, Bryon Aragam, Bernhard Schölkopf et al.NeurIPS 2024 · 37 citations
- Uncovering Meanings of Embeddings via Partial OrthogonalityYibo Jiang, Bryon Aragam, Victor VeitchNeurIPS 2023 · 20 citations
- Parts of Speech-Grounded Subspaces in Vision-Language ModelsJames Oldfield, Christos Tzelepis, Yannis Panagakis, Mihalis Nicolaou et al.NeurIPS 2023 · 13 citations
Builds on5
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation LearningWeixin Liang, Yuhui Zhang, Yongchan Kwon, Serena Yeung et al.NeurIPS 2022 · 834 citations
- Learning to Compose Soft Prompts for Compositional Zero-Shot LearningNihal V. Nayak, Peilin Yu, Stephen H. BachICLR 2023 · 41 citations
- BERT is to NLP what AlexNet is to CV: Can Pre-Trained Language Models Identify Analogies?Asahi Ushio, Luis Espinosa Anke, Steven Schockaert, José Camacho-ColladosACL 2021
Related papers
- Not Only Text: Exploring Compositionality of Visual Representations in Vision-Language ModelsDavide Berasi, Matteo Farina, Massimiliano Mancini, Elisa Ricci et al.CVPR 2025
- WRING Out The Bias: A Rotation-Based Alternative To Projection DebiasingWalter Gerych, Cassandra Parent, Quinn Perian, Rafiya Javed et al.ICLR 2026
- Cross-Modal Masked Compositional Concept Modeling for Enhancing Visio-Linguistic CompositionalityWei Li, Zhen Huang, Xinmei TianACL 2026
- Object-centric binding in Contrastive Language-Image PretrainingRim Assouel, Pietro Astolfi, Florian Bordes, Michal Drozdzal et al.NeurIPS 2025 · 14 citations
- Joint Vision-Language Social Bias Removal for CLIPHaoyu Zhang, Yangyang Guo, Mohan S. KankanhalliCVPR 2025
