Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference
Hao Yin, Guangzong Si, Zilei Wang
摘要
Multimodal large language models (MLLMs) improve performance on vision-language tasks by integrating visual features from pre-trained vision encoders into large language models (LLMs). However, how MLLMs process and utilize visual information remains unclear. In this paper, a shift in the dominant flow of visual information is uncovered: (1) in shallow layers, strong interactions are observed between image tokens and instruction tokens, where most visual information is injected into instruction tokens to form cross-modal semantic representations; (2) in deeper layers, image tokens primarily interact with each other, aggregating the remaining visual information to optimize semantic representations within visual modality. Based on these insights, we propose Hierarchical Modality-Aware Pruning (HiMAP), a plug-and-play inference acceleration method that dynamically prunes image tokens at specific layers, reducing computational costs by approximately 65% without sacrificing performance. Our findings offer a new understanding of visual information processing in MLLMs and provide a state-ofthe-art solution for efficient inference. The code is available at https://github.com/ustc-hyin/HiMAP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-ThoughtZihui Cheng, Qiguang Chen, Xiao Xu, Jiaqi Wang 等NeurIPS 2025 · 被引用 38 次
- Nüwa: Mending the Spatial Integrity Torn by VLM Token PruningYihong Huang, Fei Ma, Yihua Shao, Jingcai Guo 等ICLR 2026 · 被引用 15 次
- Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided AttentionJianfei Zhao, Feng Zhang, Xin Sun, Chong Feng 等CVPR 2026 · 被引用 6 次
- Sparrow: Text-Anchored Window Attention with Visual-Semantic Glimpsing for Speculative Decoding in Video LLMsLibo Zhang, Zhaoning Zhang, Wangyang Hong, Dongsheng LiACL 2026 · 被引用 3 次
- Aligning What Vision-Language Models See and Perceive with Adaptive Information FlowChengxin Liu, Wonseok Choi, Chenshuang Zhang, Tae-Hyun OhCVPR 2026 · 被引用 2 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMsYingqi Fan, Anhao Zhao, Jinlan Fu, Junlong Tong 等EMNLP 2025 · 被引用 11 次
- DCP: Dual-Cue Pruning for Efficient Large Vision-Language ModelsLei Jiang, Zixun Zhang, Yuting Zeng, Chunzhao Xie 等EMNLP 2025 · 被引用 2 次
- ST3: Accelerating Multimodal Large Language Model by Spatial-Temporal Visual Token TrimmingJiedong Zhuang, Lu Lu, Ming Dai, Rui Hu 等AAAI 2025 · 被引用 1 次
- HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early ExitHao Wu, Yingqi Fan, Dai Jinyang, Junlong Tong 等ICLR 2026 · 被引用 21 次
- Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token SelectionZixun Sun, Yubo Dong, Hehe Fan, Yi YangCVPR 2026
