It's a (Blind) Match! Towards Vision-Language Correspondence without Parallel Data
Dominik Schnaus, Nikita Araslanov, Daniel Cremers
Abstract
The platonic representation hypothesis suggests that vision and language embeddings become more homogeneous as model and dataset sizes increase. In particular, pairwise distances within each modality become more similar. This suggests that as foundation models mature, it may become possible to match vision and language embeddings in a fully unsupervised fashion, i.e. without parallel data. We present the first feasibility study, and investigate conformity of existing vision and language foundation models in the context of unsupervised, or "blind", matching. First, we formulate unsupervised matching as a quadratic assignment problem and introduce a novel heuristic that outperforms previous solvers. We also develop a technique to find optimal matching problems, for which a non-trivial match is very likely. Second, we conduct an extensive study deploying a range of vision and language models on four datasets. Our analysis reveals that for many problem instances, vision and language representations can be indeed matched without supervision. This finding opens up the exciting possibility of embedding semantic knowledge into other modalities virtually annotation-free. As a proof of concept, we showcase an unsupervised classifier, which achieves non-trivial classification accuracy without any image-text annotation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3baae953-95f7-4f40-a32b-b0b6b44543dbCited by top-tier papers8
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 69 citations
- Dynamic Reflections: Probing Video Representations with Text AlignmentMaks Ovsjanikov, Viorica Patraucean, Leonidas J. Guibas, Tyler Zhu et al.ICLR 2026 · 5 citations
- The Mechanistic Emergence of Symbol Grounding in Language ModelsShuyu Wu, Ziqiao Ma, Xiaoxi Luo, Yidong Huang et al.ICML 2026 · 4 citations
- Unique Lives, Shared World: Learning from Single-Life VideosTengda Han, Sayna Ebrahimi, Dilara Gokay, Li Yang Ku et al.CVPR 2026 · 2 citations
- Scene-Centric Unsupervised Video Panoptic SegmentationChristoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixe et al.CVPR 2026 · 1 citation
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- SOTAlign: Semi-Supervised Alignment of Unimodal Vision and Language Models via Optimal TransportSimon Roschmann, Paul KRZAKALA, Sonia Mazelet, Quentin Bouniot et al.ICML 2026 · 1 citation
- Do Vision and Language Encoders Represent the World Similarly?Mayug Maniparambil, Raiymbek Akshulakov, Yasser Abdelaziz Dahou Djilali, Mohamed El Amine Seddik et al.CVPR 2024 · 3 citations
- Let Go of Your Labels with Unsupervised TransferArtyom Gadetsky, Yulun Jiang, Maria BrbicICML 2024 · 16 citations
- The Indra Representation Hypothesis for Multimodal AlignmentJianglin Lu, Hailing Wang, Kuo Yang, Yitian Zhang et al.NeurIPS 2025 · 8 citations
- Universal 3D Shape Matching via Coarse-to-Fine Language GuidanceQinfeng Xiao, Guofeng Mei, Bo Yang, Zhang Liying et al.CVPR 2026 · 1 citation
