Auditing Prompt Caching in Language Model APIs
Chenchen Gu, Xiang Lisa Li, Rohith Kuditipudi, Percy Liang, Tatsunori Hashimoto
摘要
Prompt caching in large language models (LLMs) results in data-dependent timing variations: cached prompts are processed faster than noncached prompts. These timing differences introduce the risk of side-channel timing attacks. For example, if the cache is shared across users, an attacker could identify cached prompts from fast API response times to learn information about other users' prompts. Because prompt caching may cause privacy leakage, transparency around the caching policies of API providers is important. To this end, we develop and conduct statistical audits to detect prompt caching in real-world LLM API providers. We detect global cache sharing across users in seven API providers, including OpenAI, resulting in potential privacy leakage about users' prompts. Timing variations due to prompt caching can also result in leakage of information about model architecture. Namely, we find evidence that OpenAI's embedding model is a decoder-only Transformer, which was previously not publicly known. 1 However, prompt caching results in data-dependent timing
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Auditing Black-Box LLM APIs with a Rank-Based Uniformity TestXiaoyuan Zhu, Yaowen Ye, Tianyi Qiu, Hanlin Zhu 等ICLR 2026 · 被引用 22 次
- From Similarity to Vulnerability: Key Collision Attack on LLM Semantic CachingZHIXIANG ZHANG, Zesen Liu, Yuchong Xie, Quanfeng Huang 等ICML 2026 · 被引用 2 次
- PADD: Prefix-based Attention Divergence Detector for LLM JailbreaksZiqun Bao, Jiaqiang Niu, Yuchen Shao, Chengcheng WanWWW 2026
- Tvcache: A Tool-Value Cache for Post-Training LLM AgentsAbhishek Vijaya Kumar, Bhaskar Kataria, Byungsoo Oh, Emaad Manzoor 等ICML 2026
- HijackKV: New Threat in Position-Independent KV Cache ReuseYichi Zhang, Zhiqi Wang, Huan Zhang, Yuchen YangUSENIX Security 2026
它引用的顶会 Paper8
- Spectre Attacks: Exploiting Speculative ExecutionPaul Kocher, Jann Horn, Anders Fogh, Daniel Genkin 等S&P 2019 · 被引用 2,435 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke 等ICML 2024 · 被引用 157 次
相关 Paper
- I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM ServingGuanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang 等NDSS 2025
- Cache Me, Catch You: Cache Related Security Threats in LLM Serving FrameworksXiangFan Wu, Lingyun Ying, Guoqiang Chen, Yacong Gu 等NDSS 2026 · 被引用 6 次
- From Length to Content: Token-Length Side-Channel Attacks on LLM API Merged OutputsSijia Li, Tianyu Cui, Miao Chen, Xinjie Lin 等USENIX Security 2026
- "I've Decided to Leak": Probing Internals Behind Prompt Leakage IntentsJianshuo Dong, Yutong Zhang, Liu Yan, Zhenyu Zhong 等EMNLP 2025 · 被引用 1 次
- When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its CountermeasuresGuanlong Wu, Taojie Wang, Yao Zhang, Zheng Zhang 等NDSS 2026 · 被引用 6 次
