CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences
Ziran Qin, Yuchen Cao, Mingbao Lin, Wen Hu, Shixuan Fan, Ke Cheng, Weiyao Lin, Jianguo Li
Abstract
Large language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have alleviated the inference burden, they often fail to allocate resources rationally across layers with different attention patterns. In this paper, we introduce Cascading and Adaptive KV cache Eviction (CAKE), a novel approach that frames KV cache eviction as a "cake-slicing problem." CAKE assesses layer-specific preferences by considering attention dynamics in both spatial and temporal dimensions, allocates rational cache size for layers accordingly, and manages memory constraints in a cascading manner. This approach enables a global view of cache allocation, adaptively distributing resources across diverse attention mechanisms while maintaining memory budgets. CAKE also employs a new eviction indicator that considers the shifting importance of tokens over time, addressing limitations in existing methods that overlook temporal dynamics. Comprehensive experiments on LongBench and NeedleBench show that CAKE maintains model performance with only 3.2% of the KV cache and consistently outperforms current baselines across various models and memory constraints, particularly in low-memory settings. Additionally, CAKE achieves over 10× speedup in decoding latency compared to full cache when processing contexts of 128K tokens with FlashAttention-2. Our code is available at https://github.com/antgroup/cakekv .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f1ebc043-fb75-479d-be4b-db2e1df1b0b0Cited by top-tier papers20
- Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMsNgoc Bui, Shubham Sharma, Simran Lamba, Saumitra Mishra et al.ICLR 2026 · 19 citations
- ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal SmoothingYongqi An, Chang Lu, Kuan Zhu, Tao Yu et al.ICLR 2026 · 11 citations
- Which Heads Matter for Reasoning? RL-Guided KV Cache CompressionWenjie Du, Li Jiang, Keda TAO, Xue Liu et al.ICML 2026 · 11 citations
- NSNQuant: A Double Normalization Approach for Calibration-Free Low-Bit Vector Quantization of KV CacheDonghyun Son, Euntae Choi, Sungjoo YooNeurIPS 2025 · 8 citations
- DefensiveKV: Taming the Fragility of KV Cache Eviction in LLM InferenceYuan Feng, Haoyu Guo, Junlin Lv, S. Kevin Zhou et al.ICLR 2026 · 7 citations
Builds on18
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo QueryYixuan Wang, Shiyu Ji, Yijun Liu, Yuzhuang Xu et al.EMNLP 2025
- MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context InferenceKunxi Li, Zhonghua Jiang, Zhouzhou Shen, Zhaode Wang et al.ACL 2025
- SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM InferenceYi Zhao, Yajuan Peng, Cam-Tu Nguyen, Zuchao Li et al.NeurIPS 2025 · 8 citations
- HitKV: Activation Frequency Knows Which Tokens Are ImportantSanle Zhao, Yujuan Tan, Jing Yu, Zhuoxin Bai et al.AAAI 2026
- Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM InferenceYuan Feng, Junlin Lv, Yukun Cao, Xike Xie et al.NeurIPS 2025 · 256 citations
