Inference-Time Hyper-Scaling with KV Cache Compression
Adrian Lancucki, Konrad Staniszewski, Piotr Nawrot, Edoardo Maria Ponti
摘要
Inference-time scaling trades efficiency for increased reasoning accuracy by generating longer or more parallel sequences. However, in Transformer LLMs, generation cost is bottlenecked by the size of the key-value (KV) cache, rather than the number of generated tokens. Hence, we explore inference-time hyper-scaling: by compressing the KV cache, we can generate more tokens within the same compute budget and further improve the accuracy of scaled inference. The success of this approach, however, hinges on the ability of compression methods to preserve accuracy even at high compression ratios. To make hyper-scaling practical, we introduce Dynamic Memory Sparsification (DMS), a novel method for sparsifying KV caches that only requires 1K training steps to achieve 8 compression, while maintaining better accuracy than training-free sparse attention. Instead of prematurely discarding cached tokens, DMS delays token eviction, implicitly merging representations and preserving critical information. We demonstrate the effectiveness of inference-time hyper-scaling with DMS on multiple families of LLMs, showing that it boosts accuracy for comparable inference latency and memory load. For instance, we enhance Qwen-R1 32B by 12.0 points on AIME 24, 8.6 on GPQA, and 9.7 on LiveCodeBench on average for an equivalent number of memory reads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Fast KV Compaction via Attention MatchingAdam Zweiger, Xinghong Fu, Han Guo, Yoon KimICML 2026 · 被引用 14 次
- Fast and Expressive Multi-Byte Prediction with Probabilistic CircuitsAndreas Grivas, Lorenzo Loconte, Emile van Krieken, Piotr Nawrot 等ICML 2026 · 被引用 9 次
- KV Cache Transform Coding for Compact Storage in LLM InferenceKonrad Staniszewski, Adrian LancuckiICLR 2026 · 被引用 9 次
- AsyncSpade: Efficient Test-Time Scaling with Asynchronous Sparse DecodingShuqing Luo, Yilin Guan, Pingzhi Li, Hanrui Wang 等ICML 2026 · 被引用 2 次
- LatentCRS: A Variational EM Framework for Bridging Semantics and Behavior in LLM-based Conversational RecommendationGuanrong Li, Kuo Tian, Jinnan Qi, Qinghan Fu 等KDD 2026 · 被引用 1 次
它引用的顶会 Paper22
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- Dynamic Memory Compression: Retrofitting LLMs for Accelerated InferencePiotr Nawrot, Adrian Lancucki, Marcin Chochowski, David Tarjan 等ICML 2024 · 被引用 106 次
- DiffKV: Differentiated Memory Management for Large Language Models with Parallel KV CompactionYanqi Zhang, Yuwei Hu, Runyuan Zhao, John C. S. Lui 等SOSP 2025
- Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM ReasoningJiwon Song, Dongwon Jo, Yulhwa Kim, Jae-Joon KimNeurIPS 2025 · 被引用 25 次
- Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning ModelsJunhyuck Kim, Ethan Ewer, Taehong Moon, Jongho Park 等ICLR 2026 · 被引用 2 次
- RocketKV: Accelerating Long-Context LLM Inference via Two-Stage KV Cache CompressionPayman Behnam, Yaosheng Fu, Ritchie Zhao, Po-An Tsai 等ICML 2025
