Vision Function Layer in Multimodal LLMs
Cheng Shi, Yizhou Yu, Sibei Yang
摘要
This study identifies that visual-related functional decoding is distributed across different decoder layers in Multimodal Large Language Models (MLLMs). Typically, each function, such as counting, grounding, or OCR recognition, narrows down to two or three layers, which we define as Vision Function Layers (VFL). Additionally, the depth and its order of different VFLs exhibits a consistent pattern across different MLLMs, which is well-aligned with human behaviors (e.g., recognition occurs first, followed by counting, and then grounding). These findings are derived from Visual Token Swapping, our novel analytical framework that modifies targeted KV cache entries to precisely elucidate layer-specific functions during decoding. Furthermore, these insights offer substantial utility in tailoring MLLMs for real-world downstream applications. For instance, when LoRA training is selectively applied to VFLs whose functions align with the training data, VFL-LoRA not only outperform full-LoRA but also prevent out-of-domain function forgetting. Moreover, by analyzing the performance differential on training data when particular VFLs are ablated, VFL-select automatically classifies data by function, enabling highly efficient data selection to directly bolster corresponding capabilities. Consequently, VFL-select surpasses human experts in data selection, and achieves 98% of full-data performance with only 20% of the original dataset. This study delivers deeper comprehension of MLLM visual processing, fostering the creation of more efficient, interpretable, and robust models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming VideoXueyang Yu, Cheng Shi, Yang Wang, Sibei YangNeurIPS 2025 · 被引用 34 次
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and PredictionWeitai Kang, Jason Kuen, Mengwei Ren, Zijun Wei 等CVPR 2026 · 被引用 5 次
- WeaveTime: Streaming from Earlier Frames into Emergent Memory in VideoLLMsYulin Zhang, Cheng Shi, Sibei YangCVPR 2026 · 被引用 1 次
- RefAny3D: 3D Asset-Referenced Diffusion Models for Image GenerationHanzhuo Huang, Qingyang Bao, Zekai Gu, Zhongshuo Du 等ICLR 2026 · 被引用 1 次
- Discovering Compositional Hallucinations in LVLMsSibei Yang, Ge Zheng, Jiajin Tang, Jiaye Qian 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- Multimodal Language Models See Better When They Look ShallowerHaoran Chen, Junyan Lin, Xinghao Chen, Yue Fan 等EMNLP 2025
- Towards Interpreting Visual Information Processing in Vision-Language ModelsClement Neo, Luke Ong, Philip Torr, Mor Geva 等ICLR 2025
- LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-SteeringJinhe Bi, Yujun Wang, Haokun Chen, Xun Xiao 等ACL 2025
- Positional Cognitive Specialization: Where Do LLMs Learn to Comprehend and Speak Your Language?Luis Frentzen Salim, Lun-Wei Ku, Hsing-Kuo Kenneth PaoAAAI 2026
- Visual Perception by Large Language Model's WeightsFeipeng Ma, Hongwei Xue, Yizhou Zhou, Guangting Wang 等NeurIPS 2024 · 被引用 24 次
