Depth Gives a False Sense of Privacy: LLM Internal States Inversion
Tian Dong, Yan Meng, Shaofeng Li, Guoxing Chen, Zhen Liu, Haojin Zhu
摘要
Large Language Models (LLMs) are increasingly integrated into daily routines, yet they raise significant privacy and safety concerns. Recent research proposes collaborative inference, which outsources the early-layer inference to ensure data locality, and introduces model safety auditing based on inner neuron patterns. Both techniques expose the LLM's Internal States (ISs), which are traditionally considered irreversible to inputs due to optimization challenges and the highly abstract representations in deep layers. In this work, we challenge this assumption by proposing four inversion attacks that significantly improve the semantic similarity and token matching rate of inverted inputs. Specifically, we first develop two white-box optimization-based attacks tailored for low-depth and high-depth ISs. These attacks avoid local minima convergence, a limitation observed in prior work, through a two-phase inversion process. Then, we extend our optimization attack under more practical black-box weight access by leveraging the transferability between the source and the derived LLMs. Additionally, we introduce a generation-based attack that treats inversion as a translation task, employing an inversion model to reconstruct inputs. Extensive evaluation of short and long prompts from medical consulting and coding assistance datasets and 6 LLMs validates the effectiveness of our inversion attacks. Notably, a 4,112-token long medical consulting prompt can be nearly perfectly inverted with 86.88 F1 token matching from the middle layer of Llama-3 model. Finally, we evaluate four practical defenses that we found cannot perfectly prevent ISs inversion and draw conclusions for future mitigation design.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- INSIDE: LLMs' Internal States Retain the Power of Hallucination DetectionChao Chen, Kai Liu, Ze Chen, Yi Gu 等ICLR 2024 · 被引用 281 次
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 被引用 200 次
相关 Paper
- Prompt Inversion Attack Against Collaborative Inference of Large Language ModelsWenjie Qu, Yuguang Zhou, Yongji Wu, Tingsong Xiao 等S&P 2025
- Hidden No More: Attacking and Defending Private Third-Party LLM InferenceRahul Krishna Thomas, Louai Zahran, Erica Choi, Akilesh Potti 等ICML 2025
- Prompt Inference Attack on Distributed Large Language Model Inference FrameworksXinjian Luo, Ting Yu, Xiaokui XiaoCCS 2025
- Text Embedding Inversion Security for Multilingual Language ModelsYiyi Chen, Heather C. Lent, Johannes BjervaACL 2024 · 被引用 10 次
- Detecting Data Poisoning in Code Generation LLMs via Black-Box, Vulnerability-Oriented ScanningShenao Yan, Shan Jin, Shimaa Ahmed, Sunpreet Singh Arora 等CCS 2026
