USENIX Security2026Top-tier venue
From Length to Content: Token-Length Side-Channel Attacks on LLM API Merged Outputs
Sijia Li, Tianyu Cui, Miao Chen, Xinjie Lin, Zheyuan Gu, Xinhao Deng, Ke Xu, Qi Li
Abstract
Large Language Models (LLMs) are increasingly accessed via remote APIs over encrypted channels. While previous studies have shown that fine-grained side channels can leak information under token-by-token processing, modern LLM services typically stream outputs in multi-token chunks, seemingly mitigating such token-level leakage. However, in this paper, we demonstrate that token aggregation does not eliminate privacy risks. Instead, it introduces a merged token-length side channel. In this channel, the sizes of encrypted response chunks inadvertently expose character-length patterns of merged token groups. To study the practical implications of this leakage, we propose PromptEcho , a passive eavesdropping attack that reconstructs semantically meaningful portions of natural-language responses from observations of merged token lengths. PromptEcho frames the reconstruction task as a constrained sequence recovery problem, leveraging semantics-aware reasoning to resolve the ambiguity caused by merged token transmission. Specifically, PromptEcho first applies semantics-aware reasoning to decompose merged-token length observations, then uses probabilistic semantic alignment with a fine-tuned language model's linguistic priors to infer the most likely underlying tokens. We evaluate PromptEcho on real-world API interaction traffic from multiple commercial LLM services, correctly inferring the topic for 42.5% of sessions and reconstructing 25.5% of responses with a cosine similarity above 0.9 on OpenAI's GPT-4o and DeepSeek-V3. These findings demonstrate that merged-token transmission remains vulnerable to token-length side-channel attacks, indicating practical privacy risks from encrypted traffic observations under the evaluated API-based LLM settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 603f3db4-84e6-4daf-8e30-3d534e29b3dcBuilds on17
- ET-BERT: A Contextualized Datagram Representation with Pre-training Transformers for Encrypted Traffic ClassificationXinjie Lin, Gang Xiong, Gaopeng Gou, Zhen Li et al.WWW 2022 · 490 citations
- k-fingerprinting: A Robust Scalable Website Fingerprinting TechniqueJamie Hayes, George DanezisUSENIX Security 2016 · 474 citations
- Realtime Robust Malicious Traffic Detection via Frequency Domain AnalysisChuanpu Fu, Qi Li, Meng Shen, Ke XuCCS 2021 · 194 citations
- DeepCorr: Strong Flow Correlation Attacks on Tor Using Deep LearningMilad Nasr, Alireza Bahramali, Amir HoumansadrCCS 2018 · 187 citations
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke et al.ICML 2024 · 157 citations
Related papers
- What Was Your Prompt? A Remote Keylogging Attack on AI AssistantsRoy Weiss, Daniel Ayzenshteyn, Guy Amit, Yisroel MirskyUSENIX Security 2024 · 27 citations
- MoEcho: Exploiting Side-Channel Attacks to Compromise User Privacy in Mixture-of-Experts LLMsRuyi Ding, Tianhong Xu, Xinyi Shen, Aidong Adam Ding et al.CCS 2025
- I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM ServingGuanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang et al.NDSS 2025
- Reconstruct Your Previous Conversations! Comprehensively Investigating Privacy Leakage Risks in Conversations with GPT ModelsJunjie Chu, Zeyang Sha, Michael Backes, Yang ZhangEMNLP 2024 · 3 citations
- Auditing Prompt Caching in Language Model APIsChenchen Gu, Xiang Lisa Li, Rohith Kuditipudi, Percy Liang et al.ICML 2025
