Vision-centric Token Compression in Large Language Model
Ling Xing, Alex Jinpeng Wang, Rui Yan, Xiangbo Shu, Jinhui Tang
Abstract
Real-world applications are stretching context windows to hundreds of thousand of tokens while Large Language Models (LLMs) swell from billions to trillions of parameters. This dual expansion send compute and memory costs skyrocketing, making token compression indispensable. We introduce Vision Centric Token Compression (Vist), a slow-fast compression framework that mirrors human reading: the fast path renders distant tokens into images, letting a frozen, lightweight vision encoder skim the low-salience context; the slow path feeds the proximal window into the LLM for fine-grained reasoning. A Probability-Informed Visual Enhancement (PVE) objective masks high-frequency tokens during training, steering the Resampler to concentrate on semantically rich regions-just as skilled reader gloss over function words. On eleven in-context learning benchmarks, Vist achieves the same accuracy with 2.3 times fewer tokens, cutting FLOPs by 16% and memory by 50%. This method delivers remarkable results, outperforming the strongest text encoder-based compression method CEPE by 7.6% on average over benchmarks like TriviaQA, NQ, PopQA, NLUI, and CLIN, setting a new standard for token efficiency in LLMs. The project is at https://github.com/CSU-JPG/VIST.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c9c838da-5c9b-4ef7-a9bd-6c1c9c5c7878Cited by top-tier papers8
- AgentOCR: Reimagining Agent History via Optical Self-CompressionLang Feng, Fuchao Yang, Feng Chen, Xin Cheng et al.ACL 2026 · 17 citations
- OmniGaze: Reward-inspired Generalizable Gaze Estimation in the WildHongyu Qu, Jianan Wei, Xiangbo Shu, Yazhou Yao et al.NeurIPS 2025 · 15 citations
- MemOCR: Layout-Aware Visual Memory for Efficient Long-Horizon ReasoningYaorui Shi, Shugui Liu, Yu Yang, Wenyu Mao et al.ICML 2026 · 13 citations
- AstraNav-Memory: Contexts Compression for Long MemoryJunjun Hu, Xinda Xue, Botao Ren, Minghua Luo et al.CVPR 2026 · 5 citations
- Learning Gaussian Mixture-distributed Prototypes for 3D Scene Graph Generation from RGB-D SequencesRongxing Ding, Hongyu Qu, Xinguang Xiang, Pengpeng Li et al.ICML 2026
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 2,932 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
Related papers
- Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models Via Adaptive Token SkippingWeili Zeng, Ziyuan Huang, Kaixiang Ji, Yichao YanICCV 2025
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe PriorYulin Li, Haokun Gui, Ziyang Fan, Junjie Wang et al.NeurIPS 2025 · 18 citations
- CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMsJingyu Lei, Gaoang Wang, Der-Horng LeeCVPR 2026 · 1 citation
- Inference Optimal VLMs Need Fewer Visual Tokens and More ParametersKevin Y. Li, Sachin Goyal, João D. Semedo, J. Zico KolterICLR 2025
- METok: Multi-Stage Event-based Token Compression for Efficient Long Video UnderstandingMengyue Wang, Shuo Chen, Kristian Kersting, Volker Tresp et al.EMNLP 2025 · 5 citations
