ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
Guangda Liu, Chengwei Li, Jieru Zhao, Chenqi Zhang, Minyi Guo
摘要
Large Language Models (LLMs) have been widely deployed in a variety of applications, and the context length is rapidly increasing to handle tasks such as long-document QA and complex logical reasoning. However, long context poses significant challenges for inference efficiency, including high memory costs of key-value (KV) cache and increased latency due to extensive memory accesses. Recent works have proposed compressing KV cache to approximate computation, but these methods either evict tokens permanently, never recalling them for later inference, or recall previous tokens at the granularity of pages divided by textual positions. Both approaches degrade the model accuracy and output quality. To achieve efficient and accurate recallable KV cache compression, we introduce ClusterKV, which recalls tokens at the granularity of semantic clusters. We design and implement efficient algorithms and systems for clustering, selection, indexing and caching. Experiment results show that ClusterKV attains negligible accuracy loss across various tasks with 32 k context lengths, using only a 1 k to 2 k KV cache budget, and achieves up to a speedup in latency and a improvement in decoding throughput. Compared to SoTA recallable KV compression methods, ClusterKV demonstrates higher model accuracy and output quality, while maintaining or exceeding inference efficiency. Our code is available at https://github.com/sjtu-zhao-lab/ClusterKV.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo 等ACL 2025 · 被引用 334 次
- Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMsKan Zhu, Tian Tang, Qinyu Xu, Zhan Jin 等ICLR 2026 · 被引用 25 次
- FreeKV: Boosting KV Cache Retrieval for Efficient LLM InferenceGuangda Liu, Chengwei Li, Zhenyu Ning, Jing Lin 等ICLR 2026 · 被引用 17 次
- SolidAttention: Low-Latency SSD-based Serving on Memory-Constrained PCsXinrui Zheng, Dongliang Wei, Jianxiang Gao, Yixin Song 等FAST 2026 · 被引用 9 次
- SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM InferenceYi Zhao, Yajuan Peng, Cam-Tu Nguyen, Zuchao Li 等NeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper13
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao 等ICLR 2024 · 被引用 2,082 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh 等NeurIPS 2024 · 被引用 1,019 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier 等ICLR 2020 · 被引用 833 次
相关 Paper
- ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM InferenceXiang Liu, Zhenheng Tang, Peijie Dong, Zeyu Li 等NeurIPS 2025 · 被引用 71 次
- ClusterAttn: KV Cache Compression under Intrinsic Attention ClusteringMinwei Zhang, Haifeng Sun, Jingyu Wang, Shaolong Li 等ACL 2025 · 被引用 5 次
- ArkVale: Efficient Generative LLM Inference with Recallable Key-Value EvictionRenze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu 等NeurIPS 2024 · 被引用 56 次
- xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector ExtractionChi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Hung-Yueh Chiang 等ICML 2026 · 被引用 3 次
- IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMsYuzhen Mao, Qitong Wang, Martin Ester, Ke LiICLR 2026 · 被引用 6 次
