CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
Yuan Feng, Junlin Lv, Haoyu Guo, Yukun Cao, S Kevin Zhou, Xike Xie
摘要
Large language models have revolutionized natural language processing but face significant challenges of high storage and runtime costs, due to the transformer architecture's reliance on self-attention, particularly the large KV cache for long-sequence inference. Recent efforts to reduce KV cache size by pruning less critical entries based on attention weights remain empirical and lack formal grounding. This paper presents a formal study on identifying critical KV cache entries by analyzing attention output perturbation. Our analysis reveals that, beyond attention weights, the value states within KV entries and pretrained parameter matrices are also crucial. Based on this, we propose a perturbation-constrained selection algorithm that optimizes the worst-case output perturbation to identify critical entries. We demonstrate that our algorithm is a universal, plug-and-play enhancement that incurs negligible computational overhead. When integrated with three state-of-the-art cache eviction methods on three distinct LLMs, our algorithm significantly reduces the compression loss by more than half on average across 29 datasets from the Ruler and LongBench benchmarks. Further perturbation analysis, at both the head and layer levels, confirms the principles underlying our effectiveness. This work offers a new, formally grounded perspective to cache eviction , opening promising avenues for future research. The code is publicly available at https://github.com/FFY0/DefensiveKV.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu 等NeurIPS 2024 · 被引用 479 次
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang 等ICLR 2024 · 被引用 432 次
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 被引用 338 次
相关 Paper
- OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM InferenceYuzhe Gu, Xiyu Liang, Jiaojiao Zhao, Enmao DiaoICML 2026 · 被引用 7 次
- DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM InferenceYuan Feng, Haoyu Guo, Junlin Lv, S. Kevin Zhou 等ICLR 2026 · 被引用 7 次
- Question Tells You Where the Answer Is: Intention-aware Long-Context KV Cache CompressionLiang Zhao, Xiaocheng Feng, Weihong Zhong, Lei Huang 等ACL 2026
- ThinK: Thinner Key Cache by Query-Driven PruningYuhui Xu, Zhanming Jie, Hanze Dong, Lei Wang 等ICLR 2025
- CaM: Cache Merging for Memory-efficient LLMs InferenceYuxin Zhang, Yuxuan Du, Gen Luo, Yunshan Zhong 等ICML 2024 · 被引用 66 次
