Language Model Inversion
John X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov, Alexander M. Rush
摘要
Language models produce a distribution over the next token; can we use this to recover the prompt tokens? We consider the problem of language model inversion and show that next-token probabilities contain a surprising amount of information about the preceding text. Often we can recover the text in cases where it is hidden from the user, motivating a method for recovering unknown prompts given only the model's current distribution output. We consider a variety of model access scenarios, and show how even without predictions for every token in the vocabulary one can recover the necessary probability vector through search. On Llama-2 7B, our inversion method reconstructs prompts with a BLEU of 59 and token-level F1 of 78 and recovers 27% of prompts exactly. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- In-Context Unlearning: Language Models as Few-Shot UnlearnersMartin Pawelczyk, Seth Neel, Himabindu LakkarajuICML 2024 · 被引用 217 次
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke 等ICML 2024 · 被引用 157 次
- Query-Based Adversarial Prompt GenerationJonathan Hayase, Ema Borevkovic, Nicholas Carlini, Florian Tramèr 等NeurIPS 2024 · 被引用 72 次
- Privacy Backdoors: Enhancing Membership Inference through Poisoning Pre-trained ModelsYuxin Wen, Leo Marchyok, Sanghyun Hong, Jonas Geiping 等NeurIPS 2024 · 被引用 39 次
- DPIC: Decoupling Prompt and Intrinsic Characteristics for LLM Generated Text DetectionXiao Yu, Yuang Qi, Kejiang Chen, Guoqiang Chen 等NeurIPS 2024 · 被引用 24 次
它引用的顶会 Paper15
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter 等USENIX Security 2016 · 被引用 2,088 次
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
相关 Paper
- Better Language Model Inversion by Compactly Representing Next-Token DistributionsMurtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren 等NeurIPS 2025 · 被引用 13 次
- Extracting Prompts by Inverting LLM OutputsCollin Zhang, John X. Morris, Vitaly ShmatikovEMNLP 2024 · 被引用 10 次
- Reverse Prompt Engineering: A Zero-Shot, Genetic Algorithm Approach to Language Model InversionHanqing Li, Diego KlabjanEMNLP 2025 · 被引用 1 次
- An Invariant Latent Space Perspective on Language Model InversionWentao Ye, Jiaqi Hu, Haobo Wang, Xinpeng Ti 等AAAI 2026 · 被引用 1 次
- Distribution Prompting: Understanding the Expressivity of Language Models Through the Next-Token Distributions They Can ProduceHaojin Wang, Zining Zhu, Freda ShiEMNLP 2025
