Dialogue Without Limits: Constant-Sized KV Caches for Extended Response in LLMs
Ravi Ghadia, Avinash Kumar, Gaurav Jain, Prashant J. Nair, Poulami Das
摘要
Autoregressive Transformers rely on Key-Value (KV) caching to accelerate inference. However, the linear growth of the KV cache with context length leads to excessive memory consumption and bandwidth constraints. Existing methods drop distant tokens or compress states in a lossy manner, sacrificing accuracy by discarding vital context or introducing bias. We propose MorphKV, an inference-time technique that maintains a constant-sized KV cache while preserving accuracy. MorphKV balances long-range dependencies and local coherence during text generation. It eliminates early-token bias while retaining high-fidelity context by adaptively ranking tokens through correlation-aware selection. Unlike heuristic retention or lossy compression, MorphKV iteratively refines the KV cache via lightweight updates guided by attention patterns of recent tokens. This approach captures inter-token correlation with greater accuracy, which is crucial for tasks like content creation and code generation. Our studies on long-response tasks show 52.9% memory savings and 18.2% higher accuracy on average compared to state-ofthe-art prior works, enabling efficient deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Early Silicon of Raptor: The First 3D-DRAM Accelerator for Generative InferencePrashant J. Nair, Ramyad Hadidi, Subramani Ganesh, Sangamesh Kodge 等ISCA 2026 · 被引用 4 次
- Learning to Evict from Key-Value CacheLuca Moschella, Laura Manduchi, Ozan SenerICML 2026 · 被引用 4 次
它引用的顶会 Paper10
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
相关 Paper
- Value-Guided KV Compression for LLMs via Approximated CUR DecompositionAyan Sengupta, Siddhant Chaudhary, Tanmoy ChakrabortyNeurIPS 2025 · 被引用 6 次
- RefreshKV: Updating Small KV Cache During Long-form GenerationFangyuan Xu, Tanya Goyal, Eunsol ChoiACL 2025 · 被引用 6 次
- KV Cache Transform Coding for Compact Storage in LLM InferenceKonrad Staniszewski, Adrian LancuckiICLR 2026 · 被引用 9 次
- Retrospective Sparse Attention for Efficient Long-Context GenerationSeonghwan Choi, Beomseok Kang, Dongwon Jo, Jae-Joon KimICLR 2026 · 被引用 4 次
- GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache EvictionXuelin Li, Xiangqi Jin, Linfeng ZhangEMNLP 2025
