Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and Reasoning
Yu Fu, Zefan Cai, Abedelkadir Asi, Wayne Xiong, Yue Dong, Wen Xiao
Abstract
Key-Value (KV) caching is a common technique to enhance the computational efficiency of Large Language Models (LLMs), but its memory overhead grows rapidly with input length. Prior work has shown that not all tokens are equally important for text generation, proposing layer-level KV cache compression to selectively retain key information. Recognizing the distinct roles of attention heads in generation, we propose HeadKV, a head-level KV cache compression method, and HeadKV-R2, which leverages a novel contextual reasoning ability estimation for compression. Our approach operates at the level of individual heads, estimating their importance for contextual QA tasks that require both retrieval and reasoning capabilities. Extensive experiments across diverse benchmarks (LongBench, LooGLE), model architectures (e.g., Llama-3-8B-Instruct, Mistral-7B-Instruct), and long-context abilities tests demonstrate that our head-level KV cache compression significantly outperforms strong baselines, particularly in low-resource settings (KV size = 64 & 128). Notably, our method retains just 1.5% of the KV cache while achieving 97% of the performance of the full KV cache on the contextual question answering benchmark. All code and data are available at https://github.com/FYYFU/HeadKV .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b478ad3c-85be-4d1e-963d-13a4772d04f2Cited by top-tier papers43
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo et al.ACL 2025 · 334 citations
- Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM InferenceYuan Feng, Junlin Lv, Yukun Cao, Xike Xie et al.NeurIPS 2025 · 256 citations
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video UnderstandingMinsoo Kim, Kyuhong Shim, Jungwook Choi, Simyung ChangNeurIPS 2025 · 62 citations
- R-KV: Redundancy-aware KV Cache Compression for Reasoning ModelsZefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo et al.NeurIPS 2025 · 50 citations
- Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM ReasoningJiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon KimNeurIPS 2025 · 25 citations
Builds on19
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- Which Heads Matter for Reasoning? RL-Guided KV Cache CompressionWenjie Du, Li Jiang, Keda TAO, Xue Liu et al.ICML 2026 · 11 citations
- DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV CompactionYanqi Zhang, Yuwei Hu, Runyuan Zhao, John C. S. Lui et al.SOSP 2025
- RazorAttention: Efficient KV Cache Compression Through Retrieval HeadsHanlin Tang, Yang Lin, Jing Lin, Qingsen Han et al.ICLR 2025
- Value-Guided KV Compression for LLMs via Approximated CUR DecompositionAyan Sengupta, Siddhant Chaudhary, Tanmoy ChakrabortyNeurIPS 2025 · 6 citations
- FreqKV: Key-Value Compression in Frequency Domain for Context Window ExtensionJushi Kai, Yixuan Wang, Boyi Zeng, Haoli Bai et al.ICLR 2026 · 7 citations
