Question Tells You Where the Answer Is: Intention-aware Long-Context KV Cache Compression
Liang Zhao, Xiaocheng Feng, Weihong Zhong, Lei Huang, Kun Zhu, Baoxin Wang, Dayong Wu, Guoping Hu, Ting Liu, Bing Qin
摘要
The increasing context window greatly extends the capabilities of large language models, but on the other hand, it incurs an unaffordable memory overhead and computational latency due to the increasing Key-Value (KV) cache size. Recent KV cache compression methods manage to reduce the cache size by dropping irrelevant KVs. However, these methods often fail to identify crucial KVs for generation while excluding others accurately, resulting in severe information loss. To address this gap, we propose IntentKV, an intention-aware KV cache eviction method that identifies and retains crucial KVs according to the attention distribution of intention, which semantically reflects the user's goal and determines which part of the context is relevant. The consistency between the semantics and attention distribution is further substantiated through meticulously designed experiments. On this basis, Inten-tKV first distinguishes intention tokens from the vanilla context tokens based on their attention distribution distances. Then, the blockwise cumulative attention is calculated via aggregating the intention token attention. Finally, blocks that acquire high cumulative attention are picked and stored in KV cache. We evaluate our method across diverse long-context tasks and models. Results demonstrate that IntentKV can effectively maintain the model performance while reducing the KV cache size from 128K to 2K, leading to a 6.3x increase in decoding speed and 7.8x enhancement in memory efficiency compared to the default setting 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
相关 Paper
- HitKV: Activation Frequency Knows Which Tokens Are ImportantSanle Zhao, Yujuan Tan, Jing Yu, Zhuoxin Bai 等AAAI 2026
- MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context InferenceKunxi Li, Zhonghua Jiang, Zhouzhou Shen, Zhaode Wang 等ACL 2025
- NACL: A General and Effective KV Cache Eviction Framework for LLM at Inference TimeYilong Chen, Guoxia Wang, Junyuan Shang, Shiyao Cui 等ACL 2024 · 被引用 8 次
- KVzip: Query-Agnostic KV Cache Compression with Context ReconstructionJang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W. Lee 等NeurIPS 2025 · 被引用 103 次
- LazyEviction: Lagged KV Eviction with Attention Pattern Observation for Efficient Long ReasoningHaoyue Zhang, Hualei Zhang, Xiaosong Ma, Jie Zhang 等ACL 2026 · 被引用 7 次
