LangBridge: Interpreting Image as a Combination of Language Embeddings
Jiaqi Liao, Yuwei Niu, Fanqing Meng, Hao Li, Changyao Tian, Yinuo Du, Yuwen Xiong, Dianqi Li, Xizhou Zhu, Li Yuan, Jifeng Dai, Yu Cheng
Abstract
Recent years have witnessed remarkable advances in Large Vision-Language Models (LVLMs), which have achieved human-level performance across various complex vision-language tasks. Following LLaVA's paradigm, mainstream LVLMs typically employ a shallow MLP for visuallanguage alignment through a two-stage training process: pretraining for cross-modal alignment followed by instruction tuning. While this approach has proven effective, the underlying mechanisms of how MLPs bridge the modality gap remain poorly understood. Although some research has explored how LLMs process transformed visual tokens, few studies have investigated the fundamental alignment mechanism. Furthermore, the MLP adapter requires retraining whenever switching LLM backbones. To address these limitations, we first investigate the working principles of MLP adapters and discover that they learn to project visual embeddings into subspaces spanned by corresponding text embeddings progressively. Based on this insight, we propose LangBridge, a novel adapter that explicitly maps visual tokens to linear combinations of LLM vocabulary embeddings. This innovative design enables pretraining-free adapter transfer across different LLMs while maintaining performance. Our experimental results demonstrate that a LangBridge adapter pre-trained on Qwen2-0.5B can be directly applied to larger models such as LLaMA3-8B or Qwen2.5-14B while maintaining competitive performance. Overall, LangBridge enables interpretable vision-language alignment by grounding visual representations in LLM vocab embedding, while its plug-and-play design ensures efficient reuse across multiple LLMs with nearly no performance degradation. See our project page at https:// curryx-001.github.io/LangBridge.github. io/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98e2ec10-ac5e-4be5-9820-e35a4eba2a49Cited by top-tier papers4
- Look-Back: Implicit Visual Re-focusing in MLLM ReasoningShuo Yang, Yuwei Niu, Yuyang Liu, Yang Ye et al.AAAI 2026 · 32 citations
- Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking ReasoningJunhao Shen, Haiteng Zhao, Yuzhe Gu, Songyang Gao et al.NeurIPS 2025 · 13 citations
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMsBenno Krojer, Perampalli Shravan Nayak, Oscar Mañas, Vaibhav Adlakha et al.ICML 2026 · 6 citations
- EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary LabelingJiafei Song, Fengwei Zhou, Jin Qu, Wenjin Jason Li et al.CVPR 2026 · 4 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- Deep Pre-Alignment for VLMsTianyu Yu, Kechen Fang, Zihao Wan, Kaidong Zhang et al.ICML 2026
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang et al.NeurIPS 2025 · 7 citations
- SEA: Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLMsYuanyang Yin, Yaqi Zhao, Yajie Zhang, Yuanxing Zhang et al.EMNLP 2025
- Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong GeneralizationJihwan Park, Taehoon Song, Sanghyeok Lee, Miso Choi et al.AAAI 2026
- Bridging Vision and Language Spaces with Assignment PredictionJungin Park, Jiyoung Lee, Kwanghoon SohnICLR 2024 · 15 citations
