BOLT: Fewer Tokens but More Performance Retention for Efficient Vision-Language Models Inference
Jiahua Bao, Siyao Cheng, Jiaxing Du, Changjiang He, Zeming Lang, Hao Zhang, Jie Liu
Abstract
Vision-Language Models (VLMs) have achieved significant advances across various downstream tasks. However, as their performance improves, the increasing number of parameters results in slower prefilling speeds and longer inference times. To overcome these limitations, we observe that most VLMs do not require a large number of image tokens for inference, we propose BOLT (Basis-Oriented Lightweight Token-Trimming), a training-free and cross-attention-free token compression method. Unlike existing approaches, BOLT addresses the challenge of insufficient visual cues in textual prompts by leveraging token internal data distributions. We categorize tokens into three types: key tokens, proxy tokens, and remaining tokens. Then, by applying basis space similarity, we merge and filter the remaining tokens with the proxy tokens to retain the most informative ones. To account for the differences in VLM architectures and model sizes, we evaluate BOLT on LLaVA-Next-Llama3 and LLaVA-1.5 (7B and 13B). Our results show that BOLT achieves state-of-the-art performance, with a 90% token compression ratio leading to a 3.3× increase in pre-filling speed and a 1.5× improvement in inference speed, outperforming other methods.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 81cc8ccb-a85b-4279-a0bd-3c0ec3868501Cited by top-tier papers2
- Twin-T & TwintVQA: A Reliable Structure–Detail Separating VLM and a Comprehensive Benchmark for Chart and Table TasksJiahua Bao, Siyao Cheng, Jiaxing Du, Qingtao Xia et al.CVPR 2026
- Stop Mixing Things Up! BISCUIT Teaches Vision-Language Models to Learn New Concepts from Images on the SpotJiahua Bao, Siyao Cheng, Jiaxing Du, Yuhang Jia et al.AAAI 2026
Related papers
- Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-DiversityZhengyao Fang, Pengyuan Lyu, Chengquan Zhang, Guangming Lu et al.ICLR 2026 · 25 citations
- iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal ModelsLianyu Hu, Liqing Gao, Fanhua Shang, Liang Wan et al.ICLR 2026 · 9 citations
- Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language ModelsWeihao Ye, Qiong Wu, Wenhao Lin, Yiyi ZhouAAAI 2025 · 99 citations
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceAditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma, Zicheng Liu et al.CVPR 2026
- EarlyTom: Early Token Compression Completes Fast Video UnderstandingHesong Wang, Xin Jin, Lu Lu, Chenhaowen Li et al.CVPR 2026 · 7 citations
