Cross-modal Information Flow in Multimodal Large Language Models
Zhi Zhang, Srishti Yadav, Fengze Han, Ekaterina Shutova
摘要
The recent advancements in auto-regressive multimodal large language models (MLLMs) have demonstrated promising progress for vision-language tasks. While there exists a variety of studies investigating the processing of linguistic information within large language models, little is currently known about the inner working mechanism of MLLMs and how linguistic and visual information interact within these models. In this study, we aim to fill this gap by examining the information flow between different modalities-language and vision-in MLLMs, focusing on visual question answering. Specifically, given an image-question pair as input, we investigate where in the model and how the visual and linguistic information are combined to generate the final prediction. Conducting experiments with a series of models from the LLaVA series, we find that there are two distinct stages in the process of integration of the two modalities. In the lower layers, the model first transfers the more general visual features of the whole image into the representations of (linguistic) question tokens. In the middle layers, it once again transfers visual information about specific objects relevant to the question to the respective token positions of the question. Finally, in the higher layers, the resulting multimodal representation is propagated to the last position of the input sequence for the final prediction. Overall, our findings provide a new and comprehensive perspective on the spatial and functional aspects of image and language processing in the MLLMs, thereby facilitating future research into multimodal information localization and editing. Our code and collected dataset are released here: https: / / github . com / FightingFighting / crossmodal-information-flow-in-MLLM.git.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper35
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMsYaniv Nikankin, Dana Arad, Yossi Gandelsman, Yonatan BelinkovNeurIPS 2025 · 被引用 37 次
- Nüwa: Mending the Spatial Integrity Torn by VLM Token PruningYihong Huang, Fei Ma, Yihua Shao, Jingcai Guo 等ICLR 2026 · 被引用 15 次
- HiconAgent: History Context-aware Policy Optimization for GUI AgentsXurui Zhou, Gongwei Chen, Yuquan Xie, Zaijing Li 等CVPR 2026 · 被引用 11 次
- IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token PruningZhichao Sun, Yidong Ma, Gang Liu, Nemo Chen 等ICLR 2026 · 被引用 11 次
- Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation ModelingYuran Wang, Bohan Zeng, Chengzhuo Tong, Wenxuan Liu 等CVPR 2026 · 被引用 9 次
它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Towards Interpreting Visual Information Processing in Vision-Language ModelsClement Neo, Luke Ong, Philip Torr, Mor Geva 等ICLR 2025
- Understanding Information Storage and Transfer in Multi-Modal Large Language ModelsSamyadeep Basu, Martin Grayson, Cecily Morrison, Besmira Nushi 等NeurIPS 2024 · 被引用 57 次
- Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question AnsweringFederico Cocchi, Nicholas Moratelli, Marcella Cornia, Lorenzo Baraldi 等CVPR 2025
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMsMinji Kim, Taekyung Kim, Bohyung HanICLR 2026 · 被引用 8 次
- VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMsYingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong 等EMNLP 2025 · 被引用 11 次
