Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
Lin Long, Changdae Oh, Seongheon Park, Sharon Li
Abstract
Large vision-language models (LVLMs) achieve strong performance on multimodal tasks, yet they often default to their language prior (LP)---memorized textual patterns from pre-training while under-utilizing visual evidence. Prior analyses of LP mostly rely on input–output probing, which fails to reveal the internal mechanisms governing when and how vision influences model behavior. To address this gap, we present the first systematic analysis of language prior through the lens of chain-of-embedding, which examines the layer-wise representation dynamics within LVLMs. Our analysis reveals a universal phenomenon: each model exhibits a Visual Integration Point (VIP), a critical layer at which visual information begins to meaningfully reshape hidden representations and influence decoding for multimodal reasoning. Building on this observation, we introduce the Total Visual Integration (TVI) estimator, which aggregates representational discrepancy beyond the VIP to quantify how strongly visual query influences response generation. Across 60 model–dataset combinations spanning 10 contemporary LVLMs and 6 benchmarks, we demonstrate that VIP consistently emerges, and that TVI reliably predicts the strength of language prior. This offers a principled toolkit for diagnosing and understanding language prior in LVLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 181b0470-2d8a-489f-9bf7-ebb2e7c26ac7Cited by top-tier papers2
- Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual IllusionsXiaoxiao Sun, Mingyang Li, Kun Yuan, Min Woo Sun et al.CVPR 2026 · 8 citations
- Visual Instruction Bottleneck TuningChangdae Oh, Jiatong Li, Shawn Im, Sharon LiNeurIPS 2025 · 7 citations
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang et al.NeurIPS 2024 · 1,029 citations
Related papers
- Question Aware Vision Transformer for Multimodal ReasoningRoy Ganz, Yair Kittenplon, Aviad Aberdam, Elad Ben-Avraham et al.CVPR 2024
- Predictive Regularization Against Visual Representation Degradation in Multimodal Large Language ModelsEnguang Wang, Qiang Wang, Yuanchen Wu, Ke Yan et al.CVPR 2026
- AlignVLM: Bridging Vision and Language Latent Spaces for Multimodal Document UnderstandingAhmed Masry, Juan A. Rodríguez, Tianyu Zhang, Suyuchen Wang et al.NeurIPS 2025 · 7 citations
- Seeing to Generalize: How Visual Data Corrects Binding ShortcutsNicolas Buzeta, Felipe del Rio, Cristian Hinostroza, Denis Parra et al.ICML 2026 · 1 citation
- VL-InterpreT: An Interactive Visualization Tool for Interpreting Vision-Language TransformersEstelle Aflalo, Meng Du, Shao-Yen Tseng, Yongfei Liu et al.CVPR 2022 · 34 citations
