Extracting Prompts by Inverting LLM Outputs
Collin Zhang, John X. Morris, Vitaly Shmatikov
2024年份
10被引次数
21顶会引用
摘要
We consider the problem of language model inversion: given outputs of a language model, we seek to extract the prompt that generated these outputs. We develop a new black-box method, output2prompt, that extracts prompts without access to the model's logits and without adversarial or jailbreaking queries. Unlike previous methods, output2prompt only needs outputs of normal user queries. To improve memory efficiency, output2prompt employs a new sparse encoding techique. We measure the efficacy of output2prompt on a variety of user and system prompts and demonstrate zero-shot transferability across different LLMs. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 被引用 69 次
- Language Models are Injective and Hence InvertibleGiorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli 等ICLR 2026 · 被引用 36 次
- Better Language Model Inversion by Compactly Representing Next-Token DistributionsMurtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren 等NeurIPS 2025 · 被引用 13 次
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot PluginsYigitcan Kaya, Anton Landerer, Stijn Pletinckx, Michelle Zimmermann 等S&P 2026 · 被引用 12 次
- DP-Fusion: Token-Level Differentially Private Inference for Large Language ModelsRushil Thareja, Preslav Nakov, Praneeth Vepakomma, Nils LukasICLR 2026 · 被引用 7 次
它引用的顶会 Paper4
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke 等ICML 2024 · 被引用 157 次
- Text Embeddings Reveal (Almost) As Much As TextJohn X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. RushEMNLP 2023 · 被引用 60 次
- Language Model InversionJohn X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov 等ICLR 2024 · 被引用 6 次
- Vec2Face: Unveil Human Faces From Their Blackbox Features in Face RecognitionChi Nhan Duong, Thanh-Dat Truong, Khoa Luu, Kha Gia Quach 等CVPR 2020
相关 Paper
- Reverse Prompt Engineering: A Zero-Shot, Genetic Algorithm Approach to Language Model InversionHanqing Li, Diego KlabjanEMNLP 2025 · 被引用 1 次
- An Invariant Latent Space Perspective on Language Model InversionWentao Ye, Jiaqi Hu, Haobo Wang, Xinpeng Ti 等AAAI 2026 · 被引用 1 次
- Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RLHao Sun, Alihan Hüyük, Mihaela van der SchaarICLR 2024 · 被引用 48 次
- PromptBoosting: Black-Box Text Classification with Ten Forward PassesBairu Hou, Joe O'Connor, Jacob Andreas, Shiyu Chang 等ICML 2023 · 被引用 56 次
- Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language ModelsAlliot Nagle, Adway Girish, Marco Bondaschi, Michael Gastpar 等NeurIPS 2024 · 被引用 23 次
