From Length to Content: Token-Length Side-Channel Attacks on LLM API Merged Outputs
Sijia Li, Tianyu Cui, Miao Chen, Xinjie Lin, Zheyuan Gu, Xinhao Deng, Ke Xu, Qi Li
摘要
Large Language Models (LLMs) are increasingly accessed via remote APIs over encrypted channels. While previous studies have shown that fine-grained side channels can leak information under token-by-token processing, modern LLM services typically stream outputs in multi-token chunks, seemingly mitigating such token-level leakage. However, in this paper, we demonstrate that token aggregation does not eliminate privacy risks. Instead, it introduces a merged token-length side channel. In this channel, the sizes of encrypted response chunks inadvertently expose character-length patterns of merged token groups. To study the practical implications of this leakage, we propose PromptEcho , a passive eavesdropping attack that reconstructs semantically meaningful portions of natural-language responses from observations of merged token lengths. PromptEcho frames the reconstruction task as a constrained sequence recovery problem, leveraging semantics-aware reasoning to resolve the ambiguity caused by merged token transmission. Specifically, PromptEcho first applies semantics-aware reasoning to decompose merged-token length observations, then uses probabilistic semantic alignment with a fine-tuned language model's linguistic priors to infer the most likely underlying tokens. We evaluate PromptEcho on real-world API interaction traffic from multiple commercial LLM services, correctly inferring the topic for 42.5% of sessions and reconstructing 25.5% of responses with a cosine similarity above 0.9 on OpenAI's GPT-4o and DeepSeek-V3. These findings demonstrate that merged-token transmission remains vulnerable to token-length side-channel attacks, indicating practical privacy risks from encrypted traffic observations under the evaluated API-based LLM settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic ClassificationXinjie Lin, Gang Xiong, Gaopeng Gou, Zhen Li 等WWW 2022 · 被引用 490 次
- k-fingerprinting: A Robust Scalable Website Fingerprinting TechniqueJamie Hayes, George DanezisUSENIX Security 2016 · 被引用 474 次
- Realtime Robust Malicious Traffic Detection via Frequency Domain AnalysisChuanpu Fu, Qi Li, Meng Shen, Ke XuCCS 2021 · 被引用 194 次
- DeepCorr: Strong Flow Correlation Attacks on Tor Using Deep LearningMilad Nasr, Alireza Bahramali, Amir HoumansadrCCS 2018 · 被引用 187 次
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke 等ICML 2024 · 被引用 157 次
相关 Paper
- What Was Your Prompt? A Remote Keylogging Attack on AI AssistantsRoy Weiss, Daniel Ayzenshteyn, Guy Amit, Yisroel MirskyUSENIX Security 2024 · 被引用 27 次
- MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMsRuyi Ding, Tianhong Xu, Xinyi Shen, Aidong Adam Ding 等CCS 2025
- I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM ServingGuanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang 等NDSS 2025
- Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT ModelsJunjie Chu, Zeyang Sha, Michael Backes, Yang ZhangEMNLP 2024 · 被引用 3 次
- Auditing Prompt Caching in Language Model APIsChenchen Gu, Xiang Lisa Li, Rohith Kuditipudi, Percy Liang 等ICML 2025
