SCoRe: Salience-Coverage Reduction for Vision Token Pruning in Vision-Language Models
Tong Xu, Hailong Shi, Xingyu Gao
Abstract
The heavy computational burden of Large Vision-Language Models (LVLMs) stems primarily from the lengthy visual token sequences generated by their vision encoders. To mitigate this, recent work has shifted towards pruning tokens within the vision encoder. However, we observe that these methods predominantly rely on a suboptimal decoupled heuristic method. This method is conceptually flawed: it is prone to sampling collapse, fails to fundamentally eliminate token redundancy, and tends to systematically discard secondary yet important semantic clusters.Addressing this limitation, this paper proposes to formalize visual token pruning as a unified Representativeness Optimization problem. We introduce SCoRe (Salience-Coverage Reduction), a unified optimization method theoretically grounded in the Weighted k-Center Problem. SCoRe constructs the final token set by greedily selecting tokens—at each iteration, choosing the token that maximizes the current set's unified representativeness score, thereby achieving the optimization of global representativeness. Extensive experiments demonstrate that SCoRe achieves State-of-the-Art (SOTA) performance across multiple benchmarks. Notably, with negligible computational overhead, our method reduces tokens by 94.4% while retaining 95% of the full performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0985b74d-fe6d-44eb-bbff-b3f5ef4d0ea7Builds on24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
Related papers
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMsJinhong Deng, Wen Li, Joey Tianyi Zhou, Yang HeNeurIPS 2025 · 23 citations
- Semantically Comprehensive Token Pruning in LVLMs via Maximizing Concept CoverageXueting Li, Qi Liu, Chenghao Xu, Xu Yang et al.ACL 2026
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning ParadigmZhenkai Wu, Xiaowen Ma, Zhenliang Ni, Dengming Zhang et al.CVPR 2026 · 6 citations
- ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative PruningWen Luo, Peng Chen, Xiaotao Huang, LiQun HuangAAAI 2026
- Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMsQizhe Zhang, Aosong Cheng, Ming Lu, Renrui Zhang et al.ICCV 2025 · 8 citations
