Auditing Prompt Caching in Language Model APIs
Chenchen Gu, Xiang Lisa Li, Rohith Kuditipudi, Percy Liang, Tatsunori Hashimoto
Abstract
Prompt caching in large language models (LLMs) results in data-dependent timing variations: cached prompts are processed faster than noncached prompts. These timing differences introduce the risk of side-channel timing attacks. For example, if the cache is shared across users, an attacker could identify cached prompts from fast API response times to learn information about other users' prompts. Because prompt caching may cause privacy leakage, transparency around the caching policies of API providers is important. To this end, we develop and conduct statistical audits to detect prompt caching in real-world LLM API providers. We detect global cache sharing across users in seven API providers, including OpenAI, resulting in potential privacy leakage about users' prompts. Timing variations due to prompt caching can also result in leakage of information about model architecture. Namely, we find evidence that OpenAI's embedding model is a decoder-only Transformer, which was previously not publicly known. 1 However, prompt caching results in data-dependent timing
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 652e8707-2690-4ced-9a27-1191ea0edfb3Cited by top-tier papers5
- Auditing Black-Box LLM APIs with a Rank-Based Uniformity TestXiaoyuan Zhu, Yaowen Ye, Tianyi Qiu, Hanlin Zhu et al.ICLR 2026 · 22 citations
- From Similarity to Vulnerability: Key Collision Attack on LLM Semantic CachingZHIXIANG ZHANG, Zesen Liu, Yuchong Xie, Quanfeng Huang et al.ICML 2026 · 2 citations
- PADD: Prefix-based Attention Divergence Detector for LLM JailbreaksZiqun Bao, Jiaqiang Niu, Yuchen Shao, Chengcheng WanWWW 2026
- Tvcache: A Tool-Value Cache for Post-Training LLM AgentsAbhishek Vijaya Kumar, Bhaskar Kataria, Byungsoo Oh, Emaad Manzoor et al.ICML 2026
- HijackKV: New Threat in Position-Independent KV Cache ReuseYichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen YangUSENIX Security 2026
Builds on8
- Spectre Attacks: Exploiting Speculative ExecutionPaul Kocher, Jann Horn, Anders Fogh, Daniel Genkin et al.S&P 2019 · 2,435 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke et al.ICML 2024 · 157 citations
Related papers
- I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM ServingGuanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang et al.NDSS 2025
- Cache Me, Catch You: Cache Related Security Threats in LLM Serving FrameworksXiangFan Wu, Lingyun Ying, Guoqiang Chen, Yacong Gu et al.NDSS 2026 · 6 citations
- From Length to Content: Token-Length Side-Channel Attacks on LLM API Merged OutputsSijia Li, Tianyu Cui, Miao Chen, Xinjie Lin et al.USENIX Security 2026
- "I've Decided to Leak": Probing Internals Behind Prompt Leakage IntentsJianshuo Dong, Yutong Zhang, Liu Yan, Zhenyu Zhong et al.EMNLP 2025 · 1 citation
- When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its CountermeasuresGuanlong Wu, Taojie Wang, Yao Zhang, Zheng Zhang et al.NDSS 2026 · 6 citations
