ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression
Guangda Liu, Chengwei Li, Jieru Zhao, Chenqi Zhang, Minyi Guo
Abstract
Large Language Models (LLMs) have been widely deployed in a variety of applications, and the context length is rapidly increasing to handle tasks such as long-document QA and complex logical reasoning. However, long context poses significant challenges for inference efficiency, including high memory costs of key-value (KV) cache and increased latency due to extensive memory accesses. Recent works have proposed compressing KV cache to approximate computation, but these methods either evict tokens permanently, never recalling them for later inference, or recall previous tokens at the granularity of pages divided by textual positions. Both approaches degrade the model accuracy and output quality. To achieve efficient and accurate recallable KV cache compression, we introduce ClusterKV, which recalls tokens at the granularity of semantic clusters. We design and implement efficient algorithms and systems for clustering, selection, indexing and caching. Experiment results show that ClusterKV attains negligible accuracy loss across various tasks with 32 k context lengths, using only a 1 k to 2 k KV cache budget, and achieves up to a speedup in latency and a improvement in decoding throughput. Compared to SoTA recallable KV compression methods, ClusterKV demonstrates higher model accuracy and output quality, while maintaining or exceeding inference efficiency. Our code is available at https://github.com/sjtu-zhao-lab/ClusterKV.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7567a0d7-4b9c-4a46-a9f0-047d3b2e857eCited by top-tier papers21
- Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse AttentionJingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo et al.ACL 2025 · 334 citations
- Tactic: Adaptive Sparse Attention with Clustering and Distribution Fitting for Long-Context LLMsKan Zhu, Tian Tang, Qinyu Xu, Zhan Jin et al.ICLR 2026 · 25 citations
- FreeKV: Boosting KV Cache Retrieval for Efficient LLM InferenceGuangda Liu, Chengwei Li, Zhenyu Ning, Jing Lin et al.ICLR 2026 · 17 citations
- SolidAttention: Low-Latency SSD-based Serving on Memory-Constrained PCsXinrui Zheng, Dongliang Wei, Jianxiang Gao, Yixin Song et al.FAST 2026 · 9 citations
- SmallKV: Small Model Assisted Compensation of KV Cache Compression for Efficient LLM InferenceYi Zhao, Yajuan Peng, Cam-Tu Nguyen, Zuchao Li et al.NeurIPS 2025 · 8 citations
Builds on13
- SWE-bench: Can Language Models Resolve Real-world Github Issues?Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao et al.ICLR 2024 · 2,082 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
Related papers
- ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM InferenceXiang Liu, Zhenheng Tang, Peijie Dong, Zeyu Li et al.NeurIPS 2025 · 71 citations
- ClusterAttn: KV Cache Compression under Intrinsic Attention ClusteringMinwei Zhang, Haifeng Sun, Jingyu Wang, Shaolong Li et al.ACL 2025 · 5 citations
- ArkVale: Efficient Generative LLM Inference with Recallable Key-Value EvictionRenze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu et al.NeurIPS 2024 · 56 citations
- xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector ExtractionChi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Hung-Yueh Chiang et al.ICML 2026 · 3 citations
- IceCache: Memory-Efficient KV-cache Management for Long-Sequence LLMsYuzhen Mao, Qitong Wang, Martin Ester, Ke LiICLR 2026 · 6 citations
