PrefixKV: Adaptive Prefix KV Cache is What Vision Instruction-Following Models Need for Efficient Generation
Ao Wang, Hui Chen, Jianchao Tan, Kefeng Zhang, Xunliang Cai, Zijia Lin, Jungong Han, Guiguang Ding
摘要
Recently, large vision-language models (LVLMs) have rapidly gained popularity for their strong generation and reasoning capabilities given diverse multimodal inputs. However, these models incur significant computational and memory overhead during inference, which greatly hinders the efficient deployment in practical scenarios. The extensive key-value (KV) cache, necessitated by the lengthy input and output sequences, notably contributes to the high inference cost. Based on this, recent works have investigated ways to reduce the KV cache size for higher efficiency. Although effective, they generally overlook the distinct importance distributions of KV vectors across layers and maintain the same cache size for each layer during the next token prediction. This results in the significant contextual information loss for certain layers, leading to notable performance decline. To address this, we present PrefixKV, where "Prefix" means the top-ranked KV based on importance rather than position in the original sequence. It reframes the challenge of determining KV cache sizes for all layers into the task of searching for the optimal global prefix configuration. With an adaptive layer-wise KV retention recipe based on binary search, the maximum contextual information can thus be preserved in each layer, facilitating the generation. Extensive experiments demonstrate that our method achieves the state-of-the-art performance compared with others. It exhibits superior inference efficiency and generation quality tradeoffs, showing promising potential for practical applications. Code is available at https://github.com/THU-MIG/PrefixKV.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FASA: FREQUENCY-AWARE SPARSE ATTENTIONYifei Wang, Yueqi Wang, Zhenrui Yue, Huimin Zeng 等ICLR 2026 · 被引用 7 次
- AirCache: Activating Inter-Modal Relevancy KV Cache Compression for Efficient Large Vision-Language Model InferenceKai Huang, Hao Zou, Bochen Wang, Ye Xi 等ICCV 2025 · 被引用 1 次
- Extending LLM Context Window with Adaptive Grouped Positional Encoding: A Training-Free MethodXinhao Xu, Jiaxin Li, Hui Chen, Zijia Lin 等ACL 2025 · 被引用 1 次
- ProtoKV: Long-context Knowledges Are Already Well-Organized Before Your QueryZhiyuan Yu, Shijian Xiao, Zhangyue Yin, Xiaoran Liu 等ICLR 2026
它引用的顶会 Paper22
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu 等ICML 2023 · 被引用 1,493 次
相关 Paper
- TGV-KV: Text-Grounded KV Eviction for Vision-Language ModelsJizhihui Liu, Ruizi Han, Miao Zhang, Rui Shao 等ICML 2026
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language ModelsXuyang Liu, Xiyan Gui, Yuchao Zhang, Linfeng ZhangICLR 2026 · 被引用 16 次
- ZipVL: Accelerating Vision-Language Models Through Dynamic Token SparsityYefei He, Feng Chen, Jing Liu, Wenqi Shao 等ICCV 2025 · 被引用 1 次
- VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference AccelerationDezhan Tu, Danylo Vashchilenko, Yuzhe Lu, Panpan XuICLR 2025
- IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model InferenceWeijian Chen, Shuibing He, Haoyang Qu, Ruidong Zhang 等FAST 2025 · 被引用 40 次
