Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM Inference
Zijie Geng, Jie Wang, Ziqi Liu, Feng Ju, Yiming Li, Xing Li, Mingxuan Yuan, Jianye Hao, Defu Lian, Enhong Chen, Feng Wu
摘要
Key-Value (KV) cache eviction—which retains the KV pairs of the most important tokens while discarding less important ones—is a critical technique for optimizing both memory usage and inference latency in large language models (LLMs). However, existing approaches often rely on simple heuristics—such as attention weights—to measure token importance, overlooking the spatial relationships be-tween token value states in the vector space. This often leads to suboptimal token selections and thus performance degradation. To tackle this problem, we propose a novel method, namely AnDPro ( An chor D irection Pro jection), which introduces a projection-based scoring function to more accurately measure token importance. Specifically, AnDPro operates in the space of value vectors and leverages the projections of these vectors onto an “Anchor Direction” —the direction of the pre-eviction output—to measure token importance and guide more accurate token selection. Experiments on 16 datasets from the LongBench benchmark demonstrate that AnDPro can maintain 96 . 07% of the full cache accuracy using only 3 . 44% KV cache budget, reducing KV cache budget size by 46 . 0% without compromising quality compared to previous state-of-the-arts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without GenerationJinwoo Ahn, Ingyu Seong, Akhil Kedia, Junhan Kim 等ICLR 2026 · 被引用 8 次
- AttentionPredictor: Temporal Patterns Matter for KV Cache CompressionQingyue Yang, Jie Wang, Xing Li, Zhihai Wang 等NeurIPS 2025 · 被引用 8 次
- LogicTree: Improving Complex Reasoning of LLMs via Instantiated Multi-step Synthetic Logical DataZehao Wang, Lin F. Yang, Jie Wang, Kehan Wang 等NeurIPS 2025 · 被引用 5 次
它引用的顶会 Paper25
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- ZeroQuant: Efficient and Affordable Post-Training Quantization for Large-Scale TransformersZhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu 等NeurIPS 2022 · 被引用 816 次
- KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache QuantizationColeman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney 等NeurIPS 2024 · 被引用 738 次
相关 Paper
- Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo QueryYixuan Wang, Shiyu Ji, Yijun Liu, Yuzhuang Xu 等EMNLP 2025
- IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM InferenceXintong Yang, Hao Gu, Binxing Xu, Lujun Li 等ICML 2026 · 被引用 2 次
- Cache What Lasts: Token Retention for Memory-Bounded KV Cache in LLMsNgoc Bui, Shubham Sharma, Simran Lamba, Saumitra Mishra 等ICLR 2026 · 被引用 19 次
- Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache EvictionZiyao Tang, Pengkun Jiao, Xinhang Chen, LiuWei Liu 等ICML 2026 · 被引用 1 次
- GraphKV: Breaking the Static Selection Paradigm with Graph-Based KV Cache EvictionXuelin Li, Xiangqi Jin, Linfeng ZhangEMNLP 2025
