AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
Qingyue Yang, Jie Wang, Xing Li, Zhihai Wang, Chen Chen, Lei Chen, Xianzhi Yu, Wulong Liu, Jianye Hao, Mingxuan Yuan, Bin Li
摘要
With the development of large language models (LLMs), efficient inference through Key-Value (KV) cache compression has attracted considerable attention, especially for long-context generation. To compress the KV cache, recent methods identify critical KV tokens through static modeling of attention scores. However, these methods often struggle to accurately determine critical tokens as they neglect the temporal patterns in attention scores, resulting in a noticeable degradation in LLM performance. To address this challenge, we propose AttentionPredictor, which is the first learning-based method to directly predict attention patterns for KV cache compression and critical token identification. Specifically, Attention-Predictor learns a lightweight, unified convolution model to dynamically capture spatiotemporal patterns and predict the next-token attention scores. An appealing feature of AttentionPredictor is that it accurately predicts the attention score and shares the unified prediction model, which consumes negligible memory, among all transformer layers. Moreover, we propose a cross-token critical cache prefetching framework that hides the token estimation time overhead to accelerate the decoding stage. By retaining most of the attention information, AttentionPredictor achieves 13× KV cache compression and 5.6× speedup in a cache offloading scenario with comparable LLM performance, significantly outperforming the stateof-the-arts. The code is available at https://github.com/MIRALab-USTC/LLM- AttentionPredictor.
- This work was done when Qingyue Yang was an intern at Huawei. ‡ Project lead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Beyond Speedup - Utilizing KV Cache for Sampling and ReasoningZeyu XING, Xing Li, Hui-Ling Zhen, Mingxuan Yuan 等ICLR 2026 · 被引用 3 次
- Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention ModelsDifan Deng, Andreas B. Winje, Lukas Fehring, Marius LindauerICML 2026
它引用的顶会 Paper32
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
相关 Paper
- Question Tells You Where the Answer Is: Intention-aware Long-Context KV Cache CompressionLiang Zhao, Xiaocheng Feng, Weihong Zhong, Lei Huang 等ACL 2026
- ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token IdentificationYefei He, Luoming Zhang, Weijia Wu, Jing Liu 等NeurIPS 2024 · 被引用 100 次
- VL-Cache: Sparsity and Modality-Aware KV Cache Compression for Vision-Language Model Inference AccelerationDezhan Tu, Danylo Vashchilenko, Yuzhe Lu, Panpan XuICLR 2025
- A Simple and Effective L_2 Norm-Based Strategy for KV Cache CompressionAlessio Devoto, Yu Zhao, Simone Scardapane, Pasquale MinerviniEMNLP 2024 · 被引用 3 次
- TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache SelectionWei Wu, Zhuoshi Pan, Kun Fu, Chao Wang 等EMNLP 2025 · 被引用 2 次
