Towards Interpreting Visual Information Processing in Vision-Language Models
Clement Neo, Luke Ong, Philip Torr, Mor Geva, David Krueger, Fazl Barez
Abstract
Vision-Language Models (VLMs) are powerful tools for processing and understanding text and images. We study the processing of visual tokens in the language model component of LLaVA, a prominent VLM. Our approach focuses on analyzing the localization of object information, the evolution of visual token representations across layers, and the mechanism of integrating visual information for predictions. Through ablation studies, we demonstrated that object identification accuracy drops by over 70% when object-specific tokens are removed. We observed that visual token representations become increasingly interpretable in the vocabulary space across layers, suggesting an alignment with textual tokens corresponding to image content. Finally, we found that the model extracts object information from these refined representations at the last token position for prediction, mirroring the process in text-only language models for factual association tasks. These findings provide crucial insights into how VLMs process and integrate visual information, bridging the gap between our understanding of language and vision models, and paving the way for more interpretable and controllable multimodal systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cda61802-b403-4d9b-986b-e5382c652f74Cited by top-tier papers63
- VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM AgentsKangrui Wang, Pingyue Zhang, Zihan Wang, Yaning Gao et al.NeurIPS 2025 · 66 citations
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video UnderstandingMinsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung ChangNeurIPS 2025 · 62 citations
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context RetentionXin Zou, Di Lu, Yizhou Wang, Yibo Yan et al.NeurIPS 2025 · 49 citations
- GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K ResolutionFengxiang Wang, Mingshuo Chen, Yueying Li, Di Wang et al.NeurIPS 2025 · 46 citations
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMsYaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan BelinkovNeurIPS 2025 · 37 citations
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- Cross-modal Information Flow in Multimodal Large Language ModelsZhi Zhang, Srishti Yadav, Fengze Han, Ekaterina ShutovaCVPR 2025
- Mechanisms of Object Localization in Vision-Language ModelsTimothy Schaumlöffel, Martina G. Vilas, Gemma RoigCVPR 2026 · 1 citation
- What's in the Image? A Deep-Dive into the Vision of Vision Language ModelsOmri Kaduri, Shai Bagon, Tali DekelCVPR 2025
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMsHaoran Lou, Chunxiao Fan, Ziyan Liu, Yuexin Wu et al.ICCV 2025 · 1 citation
- The Narrow Gate: Localized Image-Text Communication in Native Multimodal ModelsAlessandro Serra, Francesco Ortu, Emanuele Panizon, Lucrezia Valeriani et al.NeurIPS 2025 · 4 citations
