Lookahead Q-Cache: Achieving More Consistent KV Cache Eviction via Pseudo Query
Yixuan Wang, Shiyu Ji, Yijun Liu, Yuzhuang Xu, Yang Xu, Qingfu Zhu, Wanxiang Che
摘要
Large language models (LLMs) rely on keyvalue cache (KV cache) to accelerate decoding by reducing redundant computations. However, the KV cache memory usage grows substantially with longer text sequences, posing challenges for efficient deployment. Existing KV cache eviction methods prune tokens using prefilling-stage attention scores, causing inconsistency with actual inference queries, especially under tight memory budgets. In this paper, we propose Lookahead Q-Cache (LAQ), a novel eviction framework that generates lowcost pseudo lookahead queries to better approximate the true decoding-stage queries. By using these lookahead queries as the observation window for importance estimation, LAQ achieves more consistent and accurate KV cache eviction aligned with real inference scenarios. Experimental results on LongBench and Needlein-a-Haystack benchmarks show that LAQ outperforms existing methods across various budget levels, achieving a 1 ∼ 4 point improvement on LongBench under limited cache budget. Moreover, LAQ is complementary to existing approaches and can be flexibly combined to yield further improvements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without GenerationJinwoo Ahn, Ingyu Seong, Akhil Kedia, Junhan Kim 等ICLR 2026 · 被引用 8 次
- Draft-based Approximate Inference for LLMsKevin Galim, Ethan Ewer, Wonjun Kang, Minjae Lee 等ICLR 2026 · 被引用 5 次
- IndexMem: Learned KV-Cache Eviction with Latent Memory for Long-Context LLM InferenceXintong Yang, Hao Gu, Binxing Xu, Lujun Li 等ICML 2026 · 被引用 2 次
- Judge Q: Trainable Queries for Optimized Information Retention in KV Cache EvictionYijun Liu, Yixuan Wang, Yuzhuang Xu, Shiyu Ji 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper14
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test TimeZichang Liu, Aditya Desai, Fangshuo Liao, Weitao Wang 等NeurIPS 2023 · 被引用 557 次
相关 Paper
- Accurate KV Cache Eviction via Anchor Direction Projection for Efficient LLM InferenceZijie Geng, Jie Wang, Ziqi Liu, Feng Ju 等NeurIPS 2025 · 被引用 6 次
- Get More with LESS: Synthesizing Recurrence with KV Cache Compression for Efficient LLM InferenceHarry Dong, Xinyu Yang, Zhenyu Zhang, Zhangyang Wang 等ICML 2024 · 被引用 84 次
- Predicting Future Utility: Global Combinatorial Optimization for Task-Agnostic KV Cache EvictionZiyao Tang, Pengkun Jiao, Xinhang Chen, LiuWei Liu 等ICML 2026 · 被引用 1 次
- Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM InferenceYuan Feng, Junlin Lv, Yukun Cao, Xike Xie 等NeurIPS 2025 · 被引用 256 次
- ReST-KV: Robust KV Cache Eviction with Layer-wise Output Reconstruction and Spatial-Temporal SmoothingYongqi An, Chang Lu, Kuan Zhu, Tao Yu 等ICLR 2026 · 被引用 11 次
