Lune

ACM MM2025Top-tier venue

BOLT: Fewer Tokens but More Performance Retention for Efficient Vision-Language Models Inference

Jiahua Bao, Siyao Cheng, Jiaxing Du, Changjiang He, Zeming Lang, Hao Zhang, Jie Liu

2025Year
2Top-tier citations

Abstract

Vision-Language Models (VLMs) have achieved significant advances across various downstream tasks. However, as their performance improves, the increasing number of parameters results in slower prefilling speeds and longer inference times. To overcome these limitations, we observe that most VLMs do not require a large number of image tokens for inference, we propose BOLT (Basis-Oriented Lightweight Token-Trimming), a training-free and cross-attention-free token compression method. Unlike existing approaches, BOLT addresses the challenge of insufficient visual cues in textual prompts by leveraging token internal data distributions. We categorize tokens into three types: key tokens, proxy tokens, and remaining tokens. Then, by applying basis space similarity, we merge and filter the remaining tokens with the proxy tokens to retain the most informative ones. To account for the differences in VLM architectures and model sizes, we evaluate BOLT on LLaVA-Next-Llama3 and LLaVA-1.5 (7B and 13B). Our results show that BOLT achieves state-of-the-art performance, with a 90% token compression ratio leading to a 3.3× increase in pre-filling speed and a 1.5× improvement in inference speed, outperforming other methods.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 81cc8ccb-a85b-4279-a0bd-3c0ec3868501

Cited by top-tier papers2

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines