PosPrune: Visual Token Pruning with Positional Bias Correction for Efficient Large Vision-Language Models
Ziyang Wang, Mengwei Li, Hao Yin, Wenhao Liu, Zilei Wang
Abstract
Large Vision-Language Models (LVLMs) enhance performance on vision-language tasks by integrating visual features from pre-trained vision encoders into large language models (LLMs). However, the large number of visual tokens introduces significant computational overhead. Existing token pruning methods either perform global selection via [CLS]-based attention in the vision encode or prune within LLM decoding layers. These approaches face two key challenges: (1) [CLS]-based attention primarily focuses on visually salient regions across the entire image, often overlooking semantically important tokens essential for reasoning; and (2) strong positional bias in the shallow decoder layers causes the model to favor later-positioned tokens, while neglecting earlier ones that may carry critical reasoning cues. To address these issues, we propose PosPrune, a training-free, two-stage visual token pruning framework. At the vision encoder, we introduce an Asymmetric Region-aware Pruning (ARP) strategy that retains more tokens in semantically rich regions while discarding more tokens from semantically less informative regions, thus preserving spatial diversity and task-relevant details. In the LLM decoding stage, we find that the positional bias in shallow layers is primarily driven by model architecture rather than task semantics. Based on this insight, we propose a novel Positional Bias Correction (PBC) mechanism to mitigate this bias. To further reduce redundancy, we apply Maximal Marginal Relevance (MMR) to select tokens that best balance textual relevance and diversity. Extensive experiments on various LVLMs and benchmarks demonstrate the general effectiveness of our approach. Notably, when applied to LLaVA-1.5-7B, PosPrune achieves a reduction of 85% in FLOPs while preserving 98.5% of the original performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8d03215-66ed-4c28-8687-e4c6f49c4995Builds on11
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- nocaps: novel object captioning at scaleHarsh Agrawal, Peter Anderson, Karan Desai, Yufei Wang et al.ICCV 2019 · 631 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- Boosting Multimodal Large Language Models with Visual Tokens Withdrawal for Rapid InferenceZhihang Lin, Mingbao Lin, Luxi Lin, Rongrong JiAAAI 2025 · 121 citations
Related papers
- Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMsQizhe Zhang, Aosong Cheng, Ming Lu, Renrui Zhang et al.ICCV 2025 · 8 citations
- Rethinking Visual Token Reduction in LVLMs Under Cross-Modal MisalignmentRui Xu, Yunke Wang, Yong Luo, Bo DuAAAI 2026 · 7 citations
- Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language ModelsWeihao Ye, Qiong Wu, Wenhao Lin, Yiyi ZhouAAAI 2025 · 99 citations
- TransPrune: Token Transition Pruning for Efficient Large Vision-Language ModelAo Li, Yuxiang Duan, Jinghui Zhang, Congbo Ma et al.CVPR 2026 · 3 citations
- DCP: Dual-Cue Pruning for Efficient Large Vision-Language ModelsLei Jiang, Zixun Zhang, Yuting Zeng, Chunzhao Xie et al.EMNLP 2025 · 2 citations
