Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM Inference
Zhifan Luo, Shuo Shao, Su Zhang, Lijing Zhou, Yuke Hu, Chenxu Zhao, Zhihao Liu, Zhan Qin
Abstract
The Key-Value (KV) cache, which stores intermediate attention computations (Key and Value pairs) to avoid redundant calculations, is a fundamental mechanism for accelerating Large Language Model (LLM) inference. However, this efficiency optimization introduces significant yet underexplored privacy risks. This paper provides the first comprehensive analysis of these vulnerabilities, demonstrating that an adversary can reconstruct sensitive user inputs directly from the KV-cache. We design and implement three distinct attack vectors: a direct Inversion Attack, a more broadly applicable and potent Collision Attack, and a semantic-based Injection Attack. These methods demonstrate the practicality and severity of KV-cache privacy leakage issues. To mitigate this, we propose KV-Cloak, a novel, lightweight, and efficient defense mechanism. KV-Cloak uses a reversible matrix-based obfuscation scheme, combined with operator fusion, to secure the KV-cache. Our extensive experiments show that KV-Cloak effectively thwarts all proposed attacks, reducing reconstruction quality to random noise. Crucially, it achieves this robust security with virtually no degradation in model accuracy and minimal performance overhead, offering a practical solution for trustworthy LLM deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 723ec97a-641a-43f5-a19d-304076934a58Cited by top-tier papers4
- From Similarity to Vulnerability: Key Collision Attack on LLM Semantic CachingZHIXIANG ZHANG, Zesen Liu, Yuchong Xie, Quanfeng Huang et al.ICML 2026 · 2 citations
- Reading Between the Lines: Towards Reliable Black-box LLM Fingerprinting via Zeroth-order Gradient EstimationShuo Shao, Yiming Li, Hongwei Yao, Yifei Chen et al.WWW 2026
- PADD: Prefix-based Attention Divergence Detector for LLM JailbreaksZiqun Bao, Jiaqiang Niu, Yuchen Shao, Chengcheng WanWWW 2026
- HijackKV: New Threat in Position-Independent KV Cache ReuseYichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen YangUSENIX Security 2026
Builds on14
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- When Efficiency Meets Safety: A Benchmark Security Analysis of KV Cache Compression in Large Language ModelsXiaoxiao Ma, Kuofeng Gao, Zeyi Lu, Wenxi Jiang et al.ACL 2026
- Cache Me, Catch You: Cache Related Security Threats in LLM Serving FrameworksXiangFan Wu, Lingyun Ying, Guoqiang Chen, Yacong Gu et al.NDSS 2026 · 6 citations
- I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM ServingGuanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang et al.NDSS 2025
- I Know What You Said: Unveiling Hardware Cache Side-Channels in Local Large Language Model InferenceZibo Gao, Junjie Hu, Feng Guo, Yixin Zhang et al.USENIX Security 2025
- Hidden No More: Attacking and Defending Private Third-Party LLM InferenceRahul Krishna Thomas, Louai Zahran, Erica Choi, Akilesh Potti et al.ICML 2025
