Which Heads Matter for Reasoning? RL-Guided KV Cache Compression
Wenjie Du, Li Jiang, Keda TAO, Xue Liu, Huan Wang
摘要
Reasoning large language models exhibit complex reasoning behaviors via extended chain-of-thought generation that are highly fragile to information loss during decoding, creating critical challenges for KV cache compression. Existing token-dropping methods directly disrupt reasoning chains by removing intermediate steps, while head-reallocation methods, designed for retrieval tasks, fail to preserve the heads essential for generative reasoning. However, no existing method can identify which attention heads genuinely maintain reasoning consistency and control generation termination. To address this, we propose RLKV, which uses reinforcement learning as a probe to discover which heads contribute to reasoning quality by directly optimizing their cache usage against actual generation outcomes. This discovery naturally leads to an efficient compression strategy: we allocate full KV cache to reasoning-critical heads while aggressively compressing others with constant-size KV cache. Experiments reveal that a fraction of heads proves essential for reasoning, enabling 20--60% cache reduction with near-lossless performance across diverse tasks and models, and up to 2.06x end-to-end speedup at 60% reduction.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OBS-Diff: Accurate Pruning For Diffusion Models in One-ShotJunhan Zhu, Hesong Wang, Mingluo Su, Zefang Wang 等ICLR 2026 · 被引用 26 次
- EarlyTom: Early Token Compression Completes Fast Video UnderstandingHesong Wang, Xin Jin, Lu Lu, Chenhaowen Li 等CVPR 2026 · 被引用 7 次
- HARD-KV: Head-Adaptive Regularization for Decoding-time KV CompressionYuxuan Yang, Feiyang Ren, Bowen Zeng, Dalin Zhang 等ICML 2026 · 被引用 1 次
- Dynamic Thinking-Token Selection for Efficient Reasoning in Large Reasoning Modelszhenyuan guo, Tong Chen, Wenlong Meng, Chen GONG 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper23
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV CacheZirui Liu, Jiayi Yuan, Hongye Jin, Shaochen (Henry) Zhong 等ICML 2024 · 被引用 436 次
相关 Paper
- DesireKV: Decoupling Sensitivity and Importance for Reasoning-Aware KV Cache CompressionPengyu Cheng, Jiacheng Wang, Tianle Chen, Bei Liu 等AAAI 2026
- Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and ReasoningYu Fu, Zefan Cai, Abedelkadir Asi, Wayne Xiong 等ICLR 2025
- R-KV: Redundancy-aware KV Cache Compression for Reasoning ModelsZefan Cai, Wen Xiao, Hanshi Sun, Cheng Luo 等NeurIPS 2025 · 被引用 50 次
- DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV CompactionYanqi Zhang, Yuwei Hu, Runyuan Zhao, John C. S. Lui 等SOSP 2025
- Lethe: Layer- and Time-Adaptive KV Cache Pruning for Reasoning-Intensive LLM ServingHui Zeng, Daming Zhao, Pengfei Yang, WenXuan Hou 等AAAI 2026 · 被引用 2 次
