xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
Chi-Chih Chang, Wei-Cheng Lin, Chien-Yu Lin, Hung-Yueh Chiang, Yash Akhauri, Xilai Dai, Huiqiang Jiang, Yucheng Li, Luis Ceze, Kai-Chiang Wu, Mohamed Abdelfattah
Abstract
Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent studies attempt to share KV-Cache across layers, but these approaches either require expensive pretraining or rely on per-token cross-layer cosine similarity that is often limited in practice. We show, via Centered Kernel Alignment (CKA), that the dominant singular vectors of KV-Cache are well aligned across layers. Motivated by this observation, we propose xKV, a post-training compression method that jointly factorizes groupedlayer KV-Cache into a shared low-rank subspace, substantially reducing KV-Cache memory. Across widely used LLMs, xKV achieves up to 8× KV-Cache compression while preserving accuracy on long-context tasks and in multi-turn settings. To further improve efficiency, we introduce Selective Reconstruction (SR) at decode time. Combined with SR, xKV achieves up to 4.23× end-to-end speedup over the full attention baseline, and surpasses notable baselines with 30% higher throughput under a similar accuracy level. Overall, xKV provides a plug-and-play approach to reduce both memory and latency for long-context LLM inference. Our code is publicly available at: https://github.com/ abdelfattah-lab/xKV .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12e57571-ec7d-45a7-aaf7-7121ffb0c616Cited by top-tier papers2
- TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode InferenceXiaojuan Tang, Fanxu Meng, Pingzhi Tang, Yuxuan Wang et al.ASPLOS 2026
- MatKV: Trading Compute for Flash Storage in LLM InferenceKunwoo Shin, Jay H. Park, Moonwook Oh, Yohan Jo et al.ICDE 2026
Builds on14
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
- MInference 1.0: Accelerating Pre-filling for Long-Context LLMs via Dynamic Sparse AttentionHuiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu et al.NeurIPS 2024 · 479 citations
- Model Tells You What to Discard: Adaptive KV Cache Compression for LLMsSuyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang et al.ICLR 2024 · 432 citations
- RepoBench: Benchmarking Repository-Level Code Auto-Completion SystemsTianyang Liu, Canwen Xu, Julian J. McAuleyICLR 2024 · 338 citations
Related papers
- ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable CompressionGuangda Liu, Chengwei Li, Jieru Zhao, Chenqi Zhang et al.DAC 2025 · 5 citations
- Latent-Condensed Transformer for Efficient Long Context ModelingZeng You, Yaofo Chen, Qiuwu Chen, Ying Sun et al.ACL 2026
- SqueezeAttention: 2D Management of KV-Cache in LLM Inference via Layer-wise Optimal BudgetZihao Wang, Bin Cui, Shaoduo GanICLR 2025
- LeanK: Learnable K Cache Channel Pruning for Efficient DecodingYike Zhang, Zhiyuan He, Huiqiang Jiang, Chengruidong Zhang et al.EMNLP 2025
- ParisKV: Fast and Drift-Robust KV-Cache Retrieval for Long-Context LLMsYanlin Qi, Xinhang Chen, Huiqiang Jiang, Qitong Wang et al.ICML 2026
