SPLIT-VLM: Salience-Guided Partitioning towards Local Coverage for Importance-Aware Token Dropping in Vision-Language Models
Seungil Lee, Gilha lee, Hyun Kim
Abstract
Large-scale vision–language models (VLMs) excel at multimodal reasoning, yet efficiency collapses when vision tokens—often orders of magnitude more than text—dominate compute and memory. Prior token-reduction strategies typically trade off salience (which is prone to position bias and incurs extra computation) against diversity (which can under-cover salient regions and is sensitive to hyperparameters). We present SPLIT, a theoretically grounded framework that jointly preserves salience and diversity while aggressively eliminating redundancy. SPLIT (i) estimates token importance via temporal shifts of hidden states across layers—eschewing attention scores and their biases; (ii) assigns adaptive region-level budgets to guarantee localized coverage; and (iii) selects tokens using a diversity score that prioritizes distinctive, non-redundant representations. Our analysis shows that adaptive budgeting yields tighter coverage guarantees than uniform allocation, and our selection rule maintains diversity without costly tuning. Empirically, SPLIT consistently outperforms state-of-the-art on image and video understanding benchmarks. On image understanding with LLaVA-1.5-7B, SPLIT preserves over 99% accuracy with 192 vision tokens and about 92.8% with only 64 tokens, demonstrating robust performance under severe token budgets. These results indicate that SPLIT delivers scalable, attention-score-free token reduction that makes multimodal reasoning substantially more efficient without sacrificing accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2260419f-5efd-460c-a267-0f99c1ee050aBuilds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceAditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma, Zicheng Liu et al.CVPR 2026
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language ModelsJintao Tong, Wenwei Jin, Pengda Qin, Anqi Li et al.NeurIPS 2025 · 31 citations
- One Layer's Trash is Another Layer's Treasure: Adaptive Layer-wise Visual Token Selection in LVLMsYongru Chen, Kai Zhang, Zeliang Zong, Yuchen Lu et al.CVPR 2026 · 1 citation
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMsJinhong Deng, Wen Li, Joey Tianyi Zhou, Yang HeNeurIPS 2025 · 23 citations
- LLaVA-Prumerge: Adaptive Token Reduction for Efficient Large Multimodal ModelsYuzhang Shang, Mu Cai, Bingxin Xu, Yong Jae Lee et al.ICCV 2025 · 37 citations
