R-KV: Redundancy-aware KV Cache Compression for Reasoning Models
Zefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo, Yikai Zhang, Ke Wan, Yucheng Li, Yeyang Zhou, Li-Wen Chang, Jiuxiang Gu, Zhen Dong, Animashree Anandkumar
Abstract
Reasoning models have demonstrated impressive performance in self-reflection and chain-of-thought reasoning. However, they often produce excessively long outputs, leading to prohibitively large key-value (KV) caches during inference. While chain-of-thought inference significantly improves performance on complex reasoning tasks, it can also lead to reasoning failures when deployed with existing KV cache compression approaches. To address this, we propose Redundancy-aware KV Cache Compression for Reasoning models (R-KV), a novel method specifically targeting redundant tokens in reasoning models. Our method preserves nearly 100% of the full KV cache performance using only 10% of the KV cache, substantially outperforming existing KV cache baselines, which reaches only 60% of the performance. Remarkably, R-KV even achieves 105% of full KV cache performance with 16% of the KV cache. This KV-cache reduction also leads to a 90% memory saving and a 6.6× throughput over standard chain-of-thought reasoning inference. Experimental results show that R-KV consistently outperforms existing KV cache compression baselines across two mathematical reasoning datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07122823-8d82-4db5-91a7-68a2cd0ce943Cited by top-tier papers5
- TriAttention: Efficient Long Reasoning with Trigonometric KV CompressionWeian Mao, Xi Lin, Wei Huang, Yuxin Xie et al.ICML 2026 · 16 citations
- FROST: Filtering Reasoning Outliers with Attention for Efficient ReasoningHaozheng Luo, Zhuolin Jiang, Md Zahid Hasan, Yan Chen et al.ICLR 2026 · 4 citations
- Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Modelszhenyuan guo, Tong Chen, Wenlong Meng, Chen GONG et al.ICML 2026 · 1 citation
- EAKV: An Entropy-Driven Adaptive KV Compression Framework for Long Video UnderstandingHengrui Hu, Jingyu Li, Juntao Liang, Guanyu Chen et al.ICML 2026
- Less Is More: Fast and Accurate Reasoning with Cross-Head Unified Sparse AttentionLijie Yang, Zhihao Zhang, Arti Jain, Shijie Cao et al.ICML 2026
Builds on12
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test TimeZichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang et al.NeurIPS 2023 · 557 citations
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV CacheZirui Liu, Jiayi Yuan, Hongye Jin, Shaochen (Henry) Zhong et al.ICML 2024 · 436 citations
Related papers
- BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model InferenceJanghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook ChoiICML 2026
- DesireKV: Decoupling Sensitivity and Importance for Reasoning-Aware KV Cache CompressionPengyu Cheng, Jiacheng Wang, Tianle Chen, Bei Liu et al.AAAI 2026
- ThinKV: Thought-Adaptive KV Cache Compression for Efficient Reasoning ModelsAkshat Ramachandran, Marina Neseem, Charbel Sakr, Rangharajan Venkatesan et al.ICLR 2026 · 19 citations
- Which Heads Matter for Reasoning? RL-Guided KV Cache CompressionWenjie Du, Li Jiang, Keda TAO, Xue Liu et al.ICML 2026 · 11 citations
- MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context ReasoningTao Zhang, Ziqian Zeng, Hao Peng, Huiping Zhuang et al.ACL 2026 · 3 citations
