CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation Perspective
Yuan Feng, Junlin Lv, Haoyu Guo, Yukun Cao, S Kevin Zhou, Xike Xie
Abstract
Large language models have revolutionized natural language processing but face significant challenges of high storage and runtime costs, due to the transformer architecture's reliance on self-attention, particularly the large KV cache for long-sequence inference. Recent efforts to reduce KV cache size by pruning less critical entries based on attention weights remain empirical and lack formal grounding. This paper presents a formal study on identifying critical KV cache entries by analyzing attention output perturbation. Our analysis reveals that, beyond attention weights, the value states within KV entries and pretrained parameter matrices are also crucial. Based on this, we propose a perturbation-constrained selection algorithm that optimizes the worst-case output perturbation to identify critical entries. We demonstrate that our algorithm is a universal, plug-and-play enhancement that incurs negligible computational overhead. When integrated with three state-of-the-art cache eviction methods on three distinct LLMs, our algorithm significantly reduces the compression loss by more than half on average across 29 datasets from the Ruler and LongBench benchmarks. Further perturbation analysis, at both the head and layer levels, confirms the principles underlying our effectiveness. This work offers a new, formally grounded perspective to cache eviction , opening promising avenues for future research. The code is publicly available at https://github.com/FFY0/DefensiveKV.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8d567c9-06a6-4011-a33a-6eac4e639243Cited by top-tier papers1
Ask how each one uses itBuilds on13
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu et al.NeurIPS 2024 · 479 citations
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang et al.ICLR 2024 · 432 citations
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 338 citations
Related papers
- OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM InferenceYuzhe Gu, Xiyu Liang, Jiaojiao Zhao, Enmao DiaoICML 2026 · 7 citations
- DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM InferenceYuan Feng, Haoyu Guo, Junlin Lv, S. Kevin Zhou et al.ICLR 2026 · 7 citations
- Question Tells You Where the Answer Is: Intention-aware Long-Context KV Cache CompressionLiang Zhao, Xiaocheng Feng, Weihong Zhong, Lei Huang et al.ACL 2026
- ThinK: Thinner Key Cache by Query-Driven PruningYuhui Xu, Zhanming Jie, Hanze Dong, Lei Wang et al.ICLR 2025
- CaM: Cache Merging for Memory-efficient LLMs InferenceYuxin Zhang, Yuxuan Du, Gen Luo, Yunshan Zhong et al.ICML 2024 · 66 citations
