Large Vision-Language Models Get Lost in Attention
Gongli Xi, Ye Tian, Mengyu Yang, Huahui Yi, Liang Lin, Xiaoshuai Hao, Kun Wang, Wendong Wang
摘要
Despite the rapid evolution of training paradigms, the decoder backbone of large vision-language models (LVLMs) remains fundamentally rooted in the residual-connection Transformer architecture. Therefore, deciphering the distinct roles of internal modules is critical for understanding model mechanics and guiding architectural optimization. While prior statistical approaches have provided valuable attribution-based insights, they often lack a unified theoretical basis. To bridge this gap, we propose a unified framework grounded in information theory and geometry to quantify the geometric and entropic nature of residual updates. Applying this unified framework reveals a fundamental functional decoupling: Attention acts as a subspace-preserving operator focused on reconfiguration, whereas FFNs serve as subspace-expanding operators driving semantic innovation. Strikingly, further experiments demonstrate that replacing learned attention weights with predefined values (e.g., Gaussian noise) yields comparable or even superior performance across a majority of datasets relative to vanilla models. These results expose severe misallocation and redundancy in current mechanisms, suggesting that state-of-the-art LVLMs effectively "get lost in attention" rather than efficiently leveraging visual context. Our code is publicly available at this link.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper42
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- A Comprehensive Information-Decomposition Analysis of Large Vision-Language ModelsLixin Xiu, Xufang Luo, Hideki NakayamaICLR 2026 · 被引用 4 次
- Latent Attention Denoising: A Training-Free Energy-Based Framework for Mitigating Hallucinations in Vision-Language ModelsZhiwen Luo, Siyu Jiang, Weilong Jiang, Kun HeACL 2026
- -Attn: Decomposed Attention for Large Vision-and-Language ModelsChia-Wen Kuo, Sijie Zhu, Fan Chen, Xiaohui Shen 等ICCV 2025 · 被引用 1 次
- Incorporating Residual and Normalization Layers into Analysis of Masked Language ModelsGoro Kobayashi, Tatsuki Kuribayashi, Sho Yokoi, Kentaro InuiEMNLP 2021 · 被引用 28 次
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language ModelsJintao Tong, Wenwei Jin, Pengda Qin, Anqi Li 等NeurIPS 2025 · 被引用 31 次
