I Know What You Said: Unveiling Hardware Cache Side-Channels in Local Large Language Model Inference
Zibo Gao, Junjie Hu, Feng Guo, Yixin Zhang, Yinglong Han, Siyuan Liu, Haiyang Li, Zhiqiang Lv
摘要
Large Language Models (LLMs) that can be deployed locally have recently gained popularity for privacy-sensitive tasks, with companies such as Meta, Google, and Intel playing significant roles in their development. However, the security of local LLMs through the lens of hardware cache side-channels remains unexplored. In this paper, we unveil novel side-channel vulnerabilities in local LLM inference: token value and token position leakage, which can expose both the victim's input and output text, thereby compromising user privacy. Specifically, we found that adversaries can infer the token values from the cache access patterns of the token embedding operation, and deduce the token positions from the timing of autoregressive decoding phases. To demonstrate the potential of these leaks, we design a novel eavesdropping attack framework targeting both open-source and proprietary LLM inference systems. The attack framework does not directly interact with the victim's LLM and can be executed without privilege. We evaluate the attack on a range of practical local LLM deployments (e.g., Llama, Falcon, and Gemma), and the results show that our attack achieves promising accuracy. The restored output and input text have an average edit distance of 5.2% and 17.3% to the ground truth, respectively. Furthermore, the reconstructed texts achieve average cosine similarity scores of 98.7% (input) and 98.0% (output).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Cache Me, Catch You: Cache Related Security Threats in LLM Serving FrameworksXiangFan Wu, Lingyun Ying, Guoqiang Chen, Yacong Gu 等NDSS 2026 · 被引用 6 次
- TDXRay: Microarchitectural Side-Channel Analysis of Intel TDX for Real-World WorkloadsTristan Hornetz, Hosein Yavarzadeh, Albert Cheu, Adrià Gascón 等S&P 2026 · 被引用 5 次
- From Length to Content: Token-Length Side-Channel Attacks on LLM API Merged OutputsSijia Li, Tianyu Cui, Miao Chen, Xinjie Lin 等USENIX Security 2026
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- LLM-Pruner: On the Structural Pruning of Large Language ModelsXinyin Ma, Gongfan Fang, Xinchao WangNeurIPS 2023 · 被引用 994 次
- Large Language Models Are Zero-Shot Time Series ForecastersNate Gruver, Marc Finzi, Shikai Qiu, Andrew Gordon WilsonNeurIPS 2023 · 被引用 898 次
相关 Paper
- I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM ServingGuanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang 等NDSS 2025
- Prompt Inference Attack on Distributed Large Language Model Inference FrameworksXinjian Luo, Ting Yu, Xiaokui XiaoCCS 2025
- Shadow in the Cache: Unveiling and Mitigating Privacy Risks of KV-cache in LLM InferenceZhifan Luo, Shuo Shao, Su Zhang, Lijing Zhou 等NDSS 2026 · 被引用 32 次
- Auditing Prompt Caching in Language Model APIsChenchen Gu, Xiang Lisa Li, Rohith Kuditipudi, Percy Liang 等ICML 2025
- Fun-tuning: Characterizing the Vulnerability of Proprietary LLMs to Optimization-Based Prompt Injection Attacks via the Fine-Tuning InterfaceAndrey Labunets, Nishit V. Pandya, Ashish Hooda, Xiaohan Fu 等S&P 2025
