Bridging Vision and Language Spaces with Assignment Prediction
Jungin Park, Jiyoung Lee, Kwanghoon Sohn
Abstract
This paper introduces VLAP, a novel approach that bridges pretrained vision models and large language models (LLMs) to make frozen LLMs understand the visual world. VLAP transforms the embedding space of pretrained vision models into the LLMs' word embedding space using a single linear layer for efficient and general-purpose visual and language understanding. Specifically, we harness well-established word embeddings to bridge two modality embedding spaces. The visual and text representations are simultaneously assigned to a set of word embeddings within pretrained LLMs by formulating the assigning procedure as an optimal transport problem. We predict the assignment of one modality from the representation of another modality data, enforcing consistent assignments for paired multimodal data. This allows vision and language representations to contain the same information, grounding the frozen LLMs' word embedding space in visual data. Moreover, a robust semantic taxonomy of LLMs can be preserved with visual data since the LLMs interpret and reason linguistic information from correlations between word embeddings. Experimental results show that VLAP achieves substantial improvements over the previous linear transformation-based approaches across a range of vision-language tasks, including image captioning, visual question answering, and cross-modal retrieval. We also demonstrate the learned visual representations hold a semantic taxonomy of LLMs, making visual semantic arithmetic possible.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2c01304-af26-4a90-a7ee-7c6a4c0726f4Cited by top-tier papers9
- ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language ModelsDuy M. H. Nguyen, Nghiem Tuong Diep, Trung Nguyen, Hoang-Bao Le et al.NeurIPS 2025 · 7 citations
- Tree-Sliced Entropy Partial TransportViet-Hoang Tran, Thanh Tran, Thanh T. Chu, Tam Le et al.NeurIPS 2025 · 3 citations
- TextME: Bridging Unseen Modalities Through Text DescriptionsSoyeon Hong, Jinchan Kim, Jaegook You, Seungtaek Choi et al.ICML 2026
- Tree-sliced Sobolev IPMViet-Hoang Tran, Thanh Q. Tran, Thanh T. Chu, Duy-Tung Pham et al.ICLR 2026
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language ModelsJianting Tang, Yubo Wang, Haoyu Cao, Linli XuICCV 2025
Builds on27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang et al.NeurIPS 2025 · 7 citations
- mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connectionsChenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang et al.EMNLP 2022 · 159 citations
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-and-Language Pre-trainingHongwei Xue, Yupan Huang, Bei Liu, Houwen Peng et al.NeurIPS 2021 · 100 citations
- Advancing Prompt Learning through an External LayerFangming Cui, Xun Yang, Chao Wu, Liang Xiao et al.ACM MM 2024 · 3 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
