Stealing part of a production language model
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, Thomas Steinke, Jonathan Hayase, A. Feder Cooper, Katherine Lee, Matthew Jagielski, Milad Nasr, Arthur Conmy, Eric Wallace, David Rolnick, Florian Tramèr
Abstract
We introduce the first model-stealing attack that extracts precise, nontrivial information from black-box production language models like OpenAI's ChatGPT or Google's PaLM-2. Specifically, our attack recovers the embedding projection layer (up to symmetries) of a transformer model, given typical API access. For under $20 USD, our attack extracts the entire projection matrix of OpenAI's Ada and Babbage language models. We thereby confirm, for the first time, that these black-box models have a hidden dimension of 1024 and 2048, respectively. We also recover the exact hidden dimension size of the gpt-3.5-turbo model, and estimate it would cost under $2,000 in queries to recover the entire projection matrix. We conclude with potential defenses and mitigations, and discuss the implications of possible future work that could extend our attack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9e31375b-5a69-48ca-b10e-2299bd5059a9Cited by top-tier papers56
- Representation Noising: A Defence Mechanism Against Harmful FinetuningDomenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze et al.NeurIPS 2024 · 107 citations
- Harnessing the Universal Geometry of EmbeddingsRishi D. Jha, Collin Zhang, Vitaly Shmatikov, John X. MorrisNeurIPS 2025 · 69 citations
- Open LLMs are Necessary for Current Private Adaptations and Outperform their Closed AlternativesVincent Hanke, Tom Blanchard, Franziska Boenisch, Iyiola E. Olatunji et al.NeurIPS 2024 · 27 citations
- Antidistillation SamplingYash Savani, Asher Trockman, Zhili Feng, Yixuan Even Xu et al.NeurIPS 2025 · 24 citations
- Model Provenance Testing for Large Language ModelsIvica Nikolic, Teodora Baluta, Prateek SaxenaNeurIPS 2025 · 20 citations
Builds on16
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- 8-bit Optimizers via Block-wise QuantizationTim Dettmers, Mike Lewis, Sam Shleifer, Luke ZettlemoyerICLR 2022 · 457 citations
- Active Retrieval Augmented GenerationZhengbao Jiang, Frank F. Xu, Luyu Gao, Zhiqing Sun et al.EMNLP 2023 · 315 citations
- Multi-Game Decision TransformersKuang-Huei Lee, Ofir Nachum, Mengjiao Yang, Lisa Lee et al.NeurIPS 2022 · 279 citations
Related papers
- Stealing the Decoding Algorithms of Language ModelsAli Naseh, Kalpesh Krishna, Mohit Iyyer, Amir HoumansadrCCS 2023 · 12 citations
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 291 citations
- Overcoming the Retrieval Barrier: Indirect Prompt Injection in the Wild for LLM SystemsHongyan Chang, Ergute Bao, Xinjian Luo, Ting YuUSENIX Security 2026 · 24 citations
- Uncovering Prompt Elements: Cloning System Prompts from Behavioral TracesYi Qian, Fei Peng, Hao Wu, Ligeng Chen et al.ASE 2025 · 1 citation
- Scalable Extraction of Training Data from Aligned, Production Language ModelsMilad Nasr, Javier Rando, Nicholas Carlini, Jonathan Hayase et al.ICLR 2025
