Better Language Model Inversion by Compactly Representing Next-Token Distributions
Murtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren, Swabha Swayamdipta
Abstract
Language model inversion seeks to recover hidden prompts using only language model outputs. This capability has implications for security and accountability in language model deployments, such as leaking private information from an API-protected language model's system message. We propose a new method -- prompt inversion from logprob sequences (PILS) -- that recovers hidden prompts by gleaning clues from the model's next-token probabilities over the course of multiple generation steps. Our method is enabled by a key insight: The vector-valued outputs of a language model occupy a low-dimensional subspace. This enables us to losslessly compress the full next-token probability distribution over multiple generation steps using a linear map, allowing more output information to be used for inversion. Our approach yields massive gains over previous state-of-the-art methods for recovering hidden prompts, achieving 2--3.5 times higher exact recovery rates across test sets, in one case increasing the recovery rate from 17% to 60%. Our method also exhibits surprisingly good generalization behavior; for instance, an inverter trained on 16 generations steps gets 5--27 points higher prompt recovery when we increase the number of steps to 32 at test time. Furthermore, we demonstrate strong performance of our method on the more challenging task of recovering hidden system messages. We also analyze the role of verbatim repetition in prompt recovery and propose a new method for cross-family model transfer for logit-based inverters. Our findings show that next-token probabilities are a considerably more vulnerable attack surface for inversion attacks than previously known.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c412c4e-30ef-44f0-8a0d-363519466519Cited by top-tier papers2
- Language Models are Injective and Hence InvertibleGiorgos Nikolaou, Tommaso Mencattini, Donato Crisostomi, Andrea Santilli et al.ICLR 2026 · 36 citations
- Every Language Model Has a Forgery-Resistant SignatureMatthew Finlayson, Xiang Ren, Swabha SwayamdiptaICLR 2026 · 4 citations
Builds on9
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- Information Leakage in Embedding ModelsCongzheng Song, Ananth RaghunathanCCS 2020 · 200 citations
- Stealing part of a production language modelNicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke et al.ICML 2024 · 157 citations
- Text Embeddings Reveal (Almost) As Much As TextJohn X. Morris, Volodymyr Kuleshov, Vitaly Shmatikov, Alexander M. RushEMNLP 2023 · 60 citations
- Closing the Curious Case of Neural Text DegenerationMatthew Finlayson, John Hewitt, Alexander Koller, Swabha Swayamdipta et al.ICLR 2024 · 31 citations
Related papers
- Language Model InversionJohn X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov et al.ICLR 2024 · 6 citations
- An Invariant Latent Space Perspective on Language Model InversionWentao Ye, Jiaqi Hu, Haobo Wang, Xinpeng Ti et al.AAAI 2026 · 1 citation
- Extracting Prompts by Inverting LLM OutputsCollin Zhang, John X. Morris, Vitaly ShmatikovEMNLP 2024 · 10 citations
- Reverse Prompt Engineering: A Zero-Shot, Genetic Algorithm Approach to Language Model InversionHanqing Li, Diego KlabjanEMNLP 2025 · 1 citation
- Prompt Inversion Attack Against Collaborative Inference of Large Language ModelsWenjie Qu, Yuguang Zhou, Yongji Wu, Tingsong Xiao et al.S&P 2025
