Efficient Cooperation-Aware Key and Value Management for LLM Inference
Qiheng Sun, Hongwei Zhang, Junxu Liu, Haocheng Xia, Jinfei Liu, Kui Ren, Haibo Hu
摘要
Key-value (KV) caching is a widely used technique for boosting performance in database and storage systems. It keeps frequently accessed data in fast storage to minimize redundant data fetching and improve throughput. This same idea has been adopted in Large Language Models (LLMs), where it avoids recomputing the key and value states of previous tokens in attention heads during autoregressive decoding, thereby greatly accelerating inference. However, the KV cache in LLMs faces a significant challenge due to the substantial memory required to store these KV pairs in the inference process. This issue arises because each attention head in the LLM stores its own KV cache for all context tokens, leading to the cache size that grows linearly with sequence length. This has spurred research into efficient management of the KV cache of LLMs.
One of the promising directions is KV cache budget allocation, with several approaches proposing head-level allocation as they recognize that different attention heads play distinct roles. However, these methods assess each head in isolation, overlooking their cooperative contributions within the model, which results in a deviation from their true impact. To address this limitation, we propose CoKV, a novel method that efficiently manages the KV cache in LLM inference by modeling the cooperation among attention heads as a cooperative game. By attributing the contribution of each head within the model in advance, CoKV can more effectively allocate the global KV cache budget in KV cache optimization techniques such as eviction and quantization. Extensive experiments demonstrate the effectiveness of CoKV on long-context benchmarks (e.g., LongBench, NIAH, and RULER) and mathematical reasoning benchmarks (e.g., GSM8K and MATH) across multiple model families, including Qwen, Llama, and Mistral.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper30
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney 等NeurIPS 2024 · 被引用 738 次
- Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test TimeZichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang 等NeurIPS 2023 · 被引用 557 次
相关 Paper
- IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMsYuzhen Mao, Qitong Wang, Martin Ester, Ke LiICLR 2026 · 被引用 6 次
- PQCache: Product Quantization-based KVCache for Long Context LLM InferenceHailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu 等SIGMOD 2025 · 被引用 13 次
- Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and ReasoningYu Fu, Zefan Cai, Abedelkadir Asi, Wayne Xiong 等ICLR 2025
- SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal BudgetZihao Wang, Bin Cui, Shaoduo GanICLR 2025
- Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache EvictionZiyao Tang, Pengkun Jiao, Xinhang Chen, LiuWei Liu 等ICML 2026 · 被引用 1 次
