VoCo-LLaMA: Towards Vision Compression with Large Language Models
Xubing Ye, Yukang Gan, Xiaoke Huang, Yixiao Ge, Yansong Tang
Abstract
Vision-Language Models (VLMs) have achieved remarkable success in various multi-modal tasks, but they are often bottlenecked by the limited context window and high computational cost of processing high-resolution image inputs and videos. Vision compression can alleviate this problem by reducing the vision token count. Previous approaches compress vision tokens with external modules and force LLMs to understand the compressed ones, leading to visual information loss. However, the LLMs' understanding paradigm of vision tokens is not fully utilised in the compression learning process. We propose VoCo-LLaMA, the first approach to compress vision tokens using LLMs. By introducing Vision Compression tokens during the vision instruction tuning phase and leveraging attention distillation, our method distill how LLMs comprehend vision tokens into their processing of VoCo tokens. VoCo-LLaMA facilitates effective vision compression and improves the computational efficiency during the inference stage. Specifically, our method can achieve a 576× compression rate while maintaining 83.7% performance. Furthermore, through continuous training using time-series compressed token sequences of video frames, VoCo-LLaMA demonstrates the ability to understand temporal correlations, outperforming previous methods on popular video question-answering benchmarks. Our approach presents a promising way to unlock the full potential of VLMs' contextual window, enabling more scalable multi-modal applications. The project page can be accessed via https://yxxxb.github.io/VoCo-LLaMA-page .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0d91f62c-5e6b-4c68-85ac-45de4fb05dadCited by top-tier papers48
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming VideoXueyang Yu, Cheng Shi, Yang Wang, Sibei YangNeurIPS 2025 · 34 citations
- Efficient Multi-modal Large Language Models via Progressive Consistency DistillationZichen Wen, Shaobo Wang, Yufa Zhou, Junyuan Zhang et al.NeurIPS 2025 · 26 citations
- GMSA: Enhancing Context Compression via Group Merging and Layer Semantic AlignmentJiwei Tang, Zhicheng Zhang, Shunlong Wu, Jingheng Ye et al.ACL 2026 · 24 citations
- Hybrid Token Compression for Vision-Language ModelsJusheng Zhang, Xiaoyang Guo, Kaitong Cai, Qinhan Lv et al.CVPR 2026 · 23 citations
- HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early ExitHao Wu, Yingqi Fan, Dai Jinyang, Junlong Tong et al.ICLR 2026 · 21 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
Related papers
- DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and InferenceAditya Kumar Singh, Hitesh Kandala, Pratik Prabhanjan Brahma, Zicheng Liu et al.CVPR 2026
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability PerspectiveLei Lei, Jie Gu, Xiaokang Ma, Chu Tang et al.ICLR 2026 · 3 citations
- Efficient Large Multi-modal Models via Visual Context CompressionJieneng Chen, Luoxin Ye, Ju He, Zhaoyang Wang et al.NeurIPS 2024 · 49 citations
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe PriorYulin Li, Haokun Gui, Ziyang Fan, Junjie Wang et al.NeurIPS 2025 · 18 citations
- One Token per Highly Selective Frame: Towards Extreme Compression for Long Video UnderstandingZheyu Zhang, Ziqi Pang, Shixing Chen, Xiang Hao et al.NeurIPS 2025 · 5 citations
