Extracting Prompts by Inverting LLM Outputs
Collin Zhang, John X. Morris, Vitaly Shmatikov
Abstract
We consider the problem of language model inversion: given outputs of a language model, we seek to extract the prompt that generated these outputs. We develop a new black-box method, output2prompt, that extracts prompts without access to the model's logits and without adversarial or jailbreaking queries. Unlike previous methods, output2prompt only needs outputs of normal user queries. To improve memory efficiency, output2prompt employs a new sparse encoding techique. We measure the efficacy of output2prompt on a variety of user and system prompts and demonstrate zero-shot transferability across different LLMs. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers21
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 69 citations
- Language Models are Injective and Hence InvertibleGiorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli et al.ICLR 2026 · 36 citations
- Better Language Model Inversion by Compactly Representing Next-Token DistributionsMurtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren et al.NeurIPS 2025 · 13 citations
- When AI Meets the Web: Prompt Injection Risks in Third-Party AI Chatbot PluginsYigitcan Kaya, Anton Landerer, Stijn Pletinckx, Michelle Zimmermann et al.S&P 2026 · 12 citations
- DP-Fusion: Token-Level Differentially Private Inference for Large Language ModelsRushil Thareja, Preslav Nakov, Praneeth Vepakomma, Nils LukasICLR 2026 · 7 citations
Builds on4
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke et al.ICML 2024 · 157 citations
- Text Embeddings Reveal (Almost) As Much As TextJohn X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. RushEMNLP 2023 · 60 citations
- Language Model InversionJohn X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov et al.ICLR 2024 · 6 citations
- Vec2Face: Unveil Human Faces From Their Blackbox Features in Face RecognitionChi Nhan Duong, Thanh-Dat Truong, Khoa Luu, Kha Gia Quach et al.CVPR 2020
Related papers
- Reverse Prompt Engineering: A Zero-Shot, Genetic Algorithm Approach to Language Model InversionHanqing Li, Diego KlabjanEMNLP 2025 · 1 citation
- An Invariant Latent Space Perspective on Language Model InversionWentao Ye, Jiaqi Hu, Haobo Wang, Xinpeng Ti et al.AAAI 2026 · 1 citation
- Query-Dependent Prompt Evaluation and Optimization with Offline Inverse RLHao Sun, Alihan Hüyük, Mihaela van der SchaarICLR 2024 · 48 citations
- PromptBoosting: Black-Box Text Classification with Ten Forward PassesBairu Hou, Joe O'Connor, Jacob Andreas, Shiyu Chang et al.ICML 2023 · 56 citations
- Fundamental Limits of Prompt Compression: A Rate-Distortion Framework for Black-Box Language ModelsAlliot Nagle, Adway Girish, Marco Bondaschi, Michael Gastpar et al.NeurIPS 2024 · 23 citations
