Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
Yuan Feng, Junlin Lv, Yukun Cao, Xike Xie, S. Kevin Zhou
Abstract
Large Language Models have excelled in various domains but face efficiency challenges due to the growing Key-Value (KV) cache required for long-sequence inference. Recent efforts aim to reduce KV cache size by evicting vast non-critical cache elements during runtime while preserving generation quality. However, these methods typically allocate compression budgets uniformly across all attention heads, ignoring the unique attention patterns of each head. In this paper, we establish a theoretical loss upper bound between pre- and post-eviction attention output, explaining the optimization target of prior cache eviction methods, while guiding the optimization of adaptive budget allocation. Base on this, we propose Ada-KV, the first head-wise adaptive budget allocation strategy. It offers plug-and-play benefits, enabling seamless integration with prior cache eviction methods. Extensive evaluations on 13 datasets from Ruler and 16 datasets from LongBench, all conducted under both question-aware and question-agnostic scenarios, demonstrate substantial quality improvements over existing methods. Our code is available at https://github.com/FFY0/AdaKV.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1c66129-408b-4e32-a52d-bb45c7ede22cCited by top-tier papers65
- Neural Attention SearchDifan Deng, Marius LindauerNeurIPS 2025 · 431 citations
- KVzip: Query-Agnostic KV Cache Compression with Context ReconstructionJang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee et al.NeurIPS 2025 · 103 citations
- R-KV: Redundancy-aware KV Cache Compression for Reasoning ModelsZefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo et al.NeurIPS 2025 · 50 citations
- Accelerating Streaming Video Large Language Models via Hierarchical Token CompressionYiyu Wang, Xuyang Liu, Xiyan Gui, Xinying Lin et al.CVPR 2026 · 40 citations
- Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMsKan Zhu, Tian Tang, Qinyu Xu, Zhan Jin et al.ICLR 2026 · 25 citations
Builds on28
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
Related papers
- Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache EvictionZiyao Tang, Pengkun Jiao, Xinhang Chen, LiuWei Liu et al.ICML 2026 · 1 citation
- ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal SmoothingYongqi An, Chang Lu, Kuan Zhu, Tao Yu et al.ICLR 2026 · 11 citations
- Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo QueryYixuan Wang, Shiyu Ji, Yijun Liu, Yuzhuang Xu et al.EMNLP 2025
- EntroKV: Entropy-Guided Dynamic Budget Allocation for KV-Cache CompressionWenhao Gao, Haoran Cao, Yueyan Li, YongGao Xiao et al.ICML 2026
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang et al.ICLR 2024 · 432 citations
