Do Vision and Language Encoders Represent the World Similarly?
Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Mohamed El Amine Seddik, Sanath Narayan, Karttikeya Mangalam, Noel E. O'Connor
Abstract
Aligned text-image encoders such as CLIP have become the de-facto model for vision-language tasks. Further-more, modality-specific encoders achieve impressive per-formances in their respective domains. This raises a cen-tral question: does an alignment exist between uni-modal vision and language encoders since they fundamentally rep-resent the same physical world? Analyzing the latent spaces structure of vision and language models on image-caption benchmarks using the Centered Kernel Alignment (CKA), we find that the representation spaces of unaligned and aligned encoders are semantically similar. In the absence of statistical similarity in aligned encoders like CLIP, we show that a possible matching of unaligned encoders exists with-out any training. We frame this as a seeded graph-matching problem exploiting the semantic similarity between graphs and propose two methods - a Fast Quadratic Assignment Problem optimization, and a novel localized CKA metric-based matching/retrieval. We demonstrate the effectiveness of this on several downstream tasks including cross-lingual, cross-domain caption matching and image classification. Code available at github.com/mayug/0-shot-llm-vision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2b1e517-a19a-4cf2-abd7-b93ae5fc4ff2Cited by top-tier papers21
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 69 citations
- With Limited Data for Multimodal Alignment, Let the STRUCTURE Guide YouFabian Gröger, Shuo Wen, Huyen Le, Maria BrbicNeurIPS 2025 · 16 citations
- The Indra Representation Hypothesis for Multimodal AlignmentJianglin Lu, Hailing Wang, Kuo Yang, Yitian Zhang et al.NeurIPS 2025 · 8 citations
- What Do Visual Tokens Really Encode? Uncovering Sparsity and Redundancy in Multimodal Large Language ModelsYingqi Fan, Junlong Tong, Anhao Zhao, Xiaoyu ShenCVPR 2026 · 6 citations
- Dynamic Reflections: Probing Video Representations with Text AlignmentMaks Ovsjanikov, Viorica Patraucean, Leonidas J. Guibas, Tyler Zhu et al.ICLR 2026 · 5 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- Harnessing Frozen Unimodal Encoders for Flexible Multimodal AlignmentMayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Sanath Narayan et al.CVPR 2025
- IsoCLIP: Decomposing CLIP Projectors for Efficient Intra-modal AlignmentSimone Magistri, Dipam Goswami, Marco Mistretta, Bartlomiej Twardowski et al.CVPR 2026 · 4 citations
- Target Bias Is All You Need: Zero-Shot Debiasing of Vision-Language Models With Bias CorpusTaeuk Jang, Hoin Jung, Xiaoqian WangICCV 2025 · 5 citations
- Mitigate the Gap: Improving Cross-Modal Alignment in CLIPSedigheh Eslami, Gerard de MeloICLR 2025 · 1 citation
- Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language ModelsShizhan Gong, Yankai Jiang, Qi Dou, Farzan FarniaICML 2025
