A-VL: Adaptive Attention for Large Vision-Language Models
Junyang Zhang, Mu Yuan, Ruiguang Zhong, Puhan Luo, Huiyou Zhan, Ningkang Zhang, Chengchen Hu, Xiang-Yang Li
Abstract
The Large Vision-Language Model (LVLM) integrates computer vision and natural language processing techniques, offering substantial application potential. However, these models demand extensive resources during inference. Adaptive attention techniques can dynamically reduce computational redundancy and thus improve efficiency. Although current adaptive attention methods significantly reduce the memory requirements of Transformer-based language models, they are not tailored for LVLMs. We observe that LVLMs generate responses from both remote image tokens and local text tokens, and different modalities have different attention patterns. This observation inspires us to manage the attention for each modality separately. Specifically, for visual input, we store the cache of potentially useful information but only compute the most critical parts. For language input, we care more about local information. Based on our observation and analysis of vision-language attention patterns, we develop A-VL, a plug-and-play adaptive attention tailored for LVLM inference. Extensive evaluations on three vision-language tasks and five datasets show the effectiveness of our designs. Our approach A-VL outperforms existing adaptive attention methods in reducing memory usage and computational load without compromising performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ec9bdb89-2c66-425c-b0b8-e18f7ebc54dfCited by top-tier papers2
- FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language ModelsTianyu Fu, Tengxuan Liu, Qinghao Han, Guohao Dai et al.ICCV 2025 · 3 citations
- Reducing Token Redundancy in LVLMs: A Systematic Review of Token Pruning MethodsHanzhang Yuan, Mengxuan Hu, Wenhao Zhang, Tianlong Wang et al.ACL 2026
Builds on10
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang et al.ICCV 2019 · 631 citations
- Scene Text Visual Question AnsweringAli Furkan Biten, Rubèn Tito, Andrés Mafla, Lluís Gómez i Bigorda et al.ICCV 2019 · 482 citations
Related papers
- VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference AccelerationDezhan Tu, Danylo Vashchilenko, Yuzhe Lu, Panpan XuICLR 2025
- Attention-aware Inference Optimizations for Large Vision-Language Models with Memory-efficient DecodingFatih Ilhan, Gaowen Liu, Ramana Kompella, Selim Tekin et al.CVPR 2026
- VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactionsAdrian Bulat, Alberto Baldrati, Ioannis Maniadis Metaxas, Yassine Ouali et al.CVPR 2026
- ZipVL: Accelerating Vision-Language Models Through Dynamic Token SparsityYefei He, Feng Chen, Jing Liu, Wenqi Shao et al.ICCV 2025 · 1 citation
- DCP: Dual-Cue Pruning for Efficient Large Vision-Language ModelsLei Jiang, Zixun Zhang, Yuting Zeng, Chunzhao Xie et al.EMNLP 2025 · 2 citations
