OBCache: Optimal Brain KV Cache Pruning for Efficient Long-Context LLM Inference
Yuzhe Gu, Xiyu Liang, Jiaojiao Zhao, Enmao Diao
Abstract
Large language models (LLMs) with extended context windows enable powerful applications but impose significant memory overhead, as caching all key-value (KV) states scales linearly with sequence length and batch size. Existing cache eviction methods address this by exploiting attention sparsity, yet they typically rank tokens heuristically using accumulated attention weights without considering their true impact on attention outputs. We propose Optimal Brain Cache (OBCache), a principled framework that formulates cache eviction as a layer-wise structured pruning problem. Building upon the Optimal Brain Damage (OBD) theory, OBCache quantifies token saliency by measuring the perturbation in attention outputs induced by pruning tokens, with closed-form scores derived for isolated keys, isolated values, and joint key-value pairs. Our scores account not only for attention weights but also for information from value states and attention outputs, thereby enhancing existing eviction strategies with output-aware signals. Experiments on LLaMA and Qwen models demonstrate that replacing the heuristic scores in existing works, which estimate token saliency across different query positions, with OBCache's output-aware scores consistently improves longcontext accuracy. Code is available at https:// github.com/DreamSoul-AI/OBCache .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 83ba4e1a-e886-4af7-909f-1920ae467515Builds on6
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM InferenceYuan Feng, Junlin Lv, Yukun Cao, Xike Xie et al.NeurIPS 2025 · 256 citations
- KVzip: Query-Agnostic KV Cache Compression with Context ReconstructionJang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee et al.NeurIPS 2025 · 103 citations
- Evaluating Open-Domain Question Answering in the Era of Large Language ModelsEhsan Kamalloo, Nouha Dziri, Charles L. A. Clarke, Davood RafieiACL 2023 · 96 citations
- LongBench: A Bilingual, Multitask Benchmark for Long Context UnderstandingYushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu et al.ACL 2024 · 94 citations
Related papers
- Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo QueryYixuan Wang, Shiyu Ji, Yijun Liu, Yuzhuang Xu et al.EMNLP 2025
- NACL: A General and Effective KV Cache Eviction Framework for LLM at Inference TimeYilong Chen, Guoxia Wang, Junyuan Shang, Shiyao Cui et al.ACL 2024 · 8 citations
- LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long ReasoningHaoyue Zhang, Hualei Zhang, Xiaosong Ma, Jie Zhang et al.ACL 2026 · 7 citations
- CriticalKV: Optimizing KV Cache Eviction from an Output Perturbation PerspectiveYuan Feng, Junlin Lv, Haoyu Guo, Yukun Cao et al.ICML 2026 · 21 citations
- Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM InferenceZijie Geng, Jie Wang, Ziqi Liu, Feng Ju et al.NeurIPS 2025 · 6 citations
