Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMs
Ngoc Bui, Shubham Sharma, Simran Lamba, Saumitra Mishra, Rex Ying
摘要
Memory and computation remain core bottlenecks in long-horizon LLM inference due to the quadratic cost of self-attention and the ever-growing key-value (KV) cache. Existing strategies for memory-bounded inference, such as quantization, offloading, or heuristic KV eviction, either incur high orchestration costs or rely on unreliable attention-based proxies of importance. We propose TRIM-KV, a novel approach that learns each token’s intrinsic importance at creation time via a lightweight retention gate. Each gate predicts a scalar retention score that decays over time, reflecting the long-term utility of the token for a specific layer and head. Tokens with low scores are evicted when the memory budget is exceeded, ensuring that the cache always contains the most critical tokens. TRIM-KV is trained efficiently through distillation from a frozen LLM combined with a capacity loss, requiring only gate fine-tuning and adding negligible inference overhead. Across mathematical reasoning (GSM8K, MATH-500, AIME24), procedural generation (LongProc), conversational long-memory benchmarks (LongMemEval), and long-context understanding (LongBenchV2 and SCBench), TRIM-KV consistently outperforms strong eviction and learnable retrieval baselines, especially in low-memory regimes. Remarkably, it even surpasses full-cache models in some settings, showing that selective retention can serve as a form of regularization, suppressing noise from uninformative tokens. Qualitative analyses further reveal that learned retention scores align with human intuition, naturally recovering heuristics such as sink tokens, sliding windows, and gist compression without explicit design. Beyond efficiency, retention scores provide insights into layer- and head-specific roles, suggesting a new path toward LLM interpretability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- EpiCache: Episodic KV Cache Management for Long-Term Conversation on Resource-Constrained EnvironmentsMinsoo Kim, Arnav Kundu, Han-Byul Kim, Richa Dixit 等ICML 2026 · 被引用 4 次
- BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model InferenceJanghyeon Kim, Minsoo Kim, Kyuhong Shim, Jungwook ChoiICML 2026
- Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language ModelsChien Van Nguyen, Ryan A. Rossi, Linh Ngo Van, Franck Dernoncourt 等ACL 2026
它引用的顶会 Paper30
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
相关 Paper
- IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM InferenceXintong Yang, Hao Gu, Binxing Xu, Lujun Li 等ICML 2026 · 被引用 2 次
- Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM InferenceZijie Geng, Jie Wang, Ziqi Liu, Feng Ju 等NeurIPS 2025 · 被引用 6 次
- LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without GenerationJinwoo Ahn, Ingyu Seong, Akhil Kedia, Junhan Kim 等ICLR 2026 · 被引用 8 次
- Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo QueryYixuan Wang, Shiyu Ji, Yijun Liu, Yuzhuang Xu 等EMNLP 2025
- Learning to Evict from Key-Value CacheLuca Moschella, Laura Manduchi, Ozan SenerICML 2026 · 被引用 4 次
