I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM Serving
Guanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang, Jianyu Niu, Ye Wu, Yinqian Zhang
Abstract
—Large Language Models (LLMs), which laid the groundwork for Artificial General Intelligence (AGI), have recently gained significant traction in academia and industry due to their disruptive applications. In order to enable scalable applications and efficient resource management, various multi-tenant LLM serving frameworks have been proposed, in which the LLM caters to the needs of multiple users simultaneously. One notable mechanism in recent works, such as SGLang and vLLM, is sharing the Key-Value (KV) cache for identical token sequences among multiple users, saving both memory and computation. This paper presents the first investigation on security risks associated with multi-tenant LLM serving. We show that the state-of-the-art mechanisms of KV cache sharing may lead to new side channel attack vectors, allowing unauthorized reconstruction of user prompts and compromising sensitive user information among mutually distrustful users. Specifically, we introduce our attack, P ROMPT P EEK , and apply it to three scenarios where the adversary, with varying degrees of prior knowledge, is capable of reverse-engineering prompts from other users. This study underscores the need for careful resource management in multi-tenant LLM serving and provides critical insights for future security enhancement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8dd3df18-ef2b-43a7-b67f-fcfa11d3351fCited by top-tier papers8
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM InferenceZhifan Luo, Shuo Shao, Su Zhang, Lijing Zhou et al.NDSS 2026 · 32 citations
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM AgentsQizheng Zhang, Michael Wornow, Kunle OlukotunNeurIPS 2025 · 27 citations
- When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its CountermeasuresGuanlong Wu, Taojie Wang, Yao Zhang, Zheng Zhang et al.NDSS 2026 · 6 citations
- Cache Me, Catch You: Cache Related Security Threats in LLM Serving FrameworksXiangFan Wu, Lingyun Ying, Guoqiang Chen, Yacong Gu et al.NDSS 2026 · 6 citations
- From Similarity to Vulnerability: Key Collision Attack on LLM Semantic CachingZHIXIANG ZHANG, Zesen Liu, Yuchong Xie, Quanfeng Huang et al.ICML 2026 · 2 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Spectre Attacks: Exploiting Speculative ExecutionPaul Kocher, Jann Horn, Anders Fogh, Daniel Genkin et al.S&P 2019 · 2,435 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- One Bit Flips, One Cloud Flops: Cross-VM Row Hammer Attacks and Privilege EscalationYuan Xiao, Xiaokuan Zhang, Yinqian Zhang, Radu TeodorescuUSENIX Security 2016 · 272 citations
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen et al.CCS 2024 · 132 citations
Related papers
- Auditing Prompt Caching in Language Model APIsChenchen Gu, Xiang Lisa Li, Rohith Kuditipudi, Percy Liang et al.ICML 2025
- I Know What You Said: Unveiling Hardware Cache Side-Channels in Local Large Language Model InferenceZibo Gao, Junjie Hu, Feng Guo, Yixin Zhang et al.USENIX Security 2025
- From Length to Content: Token-Length Side-Channel Attacks on LLM API Merged OutputsSijia Li, Tianyu Cui, Miao Chen, Xinjie Lin et al.USENIX Security 2026
- HijackKV: New Threat in Position-Independent KV Cache ReuseYichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen YangUSENIX Security 2026
- Exploiting the Shadows: Unveiling Privacy Leaks through Lower-Ranked Tokens in Large Language ModelsYuan Zhou, Zhuo Zhang, Xiangyu ZhangACL 2025 · 2 citations
