Scalable Extraction of Training Data from Aligned, Production Language Models
Milad Nasr, Javier Rando, Nicholas Carlini, Jonathan Hayase, Matthew Jagielski, A. Feder Cooper, Daphne Ippolito, Christopher A. Choquette-Choo, Florian Tramèr, Katherine Lee
摘要
Large language models are prone to memorizing some of their training data. Memorized (and possibly sensitive) samples can then be extracted at generation time by adversarial or benign users. There is hope that model alignment-a standard training process that tunes a model to harmlessly follow user instructions-would mitigate the risk of extraction. However, we develop two novel attacks that undo a language model's alignment and recover thousands of training examples from popular proprietary aligned models such as OpenAI's ChatGPT. Our work highlights the limitations of existing safeguards to prevent training data leakage in production language models. * Equal contribution 1 While limited information is available about proprietary production models, some aligned models like GPT-4 have been trained to "refuse to answer certain types of requests," including those related to training data extraction (OpenAI, 2023). Published as a conference paper at ICLR 2025 abusing a production system's finetuning interface (Peng et al., 2023) , which allows users to further train a model on provided data (Qi et al., 2023) . These prior attacks have been successful at making aligned models output harmful content, but not training data. In this paper, we develop new attack techniques for this purpose. Note that we deliberately refrain from giving a formal definition of alignment. The primary reason is that different organizations have different definitions for what it means for a model to be "aligned with human preferences." Further, even if we have a loose or informal sense of what a particular organization views as alignment-through, e.g., high-level technical reports (OpenAI, 2023)-this does not directly translate to understanding exactly how these organizations employ concrete training techniques to align their proprietary models. We instead characterize aligned models in terms of specific behaviors they should not exhibit-in our case exact regurgitation of training data. EXPERIMENTAL SETUP Validating memorization. Typically, we would validate training data extraction by searching for the extracted text in the training dataset. However, proprietary language models like ChatGPT and Gemini do not have public training datasets. Since it is widely known that a large fraction of these models' training data is scraped from the public web, prior works have resorted to manual Google searches to check for the presence of model generations online (Carlini et al., 2021) . This is timeconsuming and does not scale. We propose a more scalable approach. First, we approximate the web-based training data of production models by building a large corpus of text from the internet-by merging (and deduplicating) four of the largest published language model training datasets: The Pile (Gao et al., 2020), RefinedWeb (Penedo et al., 2023) , RedPajama (Together, 2023a), and Dolma (Soldaini, 2023) (Appendix A.4). This corpus, which we call AUXDATASET, is the largest public index of LLM training data to date (9 terabytes). We then approximate an internet-wide search by performing a local search over this corpus. We implement a suffix array for efficient search over AUXDATASET. (See Appendix A.5 and Lee et al. (2022) for details.) We thus call a subsequence of a generation memorized if (as in Definition 1) a 50-token-length subsequence exactly appears in AUXDATASET. This validation method will only be able to provide a loose lower bound for memorization; it will undercount the success of training data extraction since AUXDATASET does not include the full training dataset for proprietary models. Moreover, we do not count sequences that are approximately memorized, i.e., a generation for which a near-exact match (e.g., a paraphrase) appears in the AUXDATASET. Nevertheless, this validation methodology satisfies our goals: we aim to provide a lower bound on the amount of extracted memorized text, and to demonstrate that we are able to extract exact subsequences of training data from aligned language models. Models. We study training data extraction in two production language model families that have been trained with alignment, ChatGPT (from OpenAI) and Gemini (from Google). For ChatGPT, we consider the two latest versions of aligned and conversational models at the moment of writing (gpt-3.5-turbo and gpt-4), and for Gemini, we consider the latest publicly available version: Gemini 1.5 Pro, a state-of-the-art model for long-context generation understanding. We compare these production, closed-weight models (i.e., embedded in systems, and which we interact with via developer APIs) with several open-weight (i.e., with publicly accessible weights), unaligned language models including GPT-2 (1.5B) (Radford et al., 2019) , GPT-Neo (6B) (Black et al., 2021) , Pythia (1.4B and 6.9B) (Biderman et al., 2023) , OPT (1.3B and 6.7B) (Zhang et al., 2022) , LLaMA (7B and 65B) (Touvron et al., 2023a), RedPajama-INCITE base (3B and 7B) (Together, 2023b), Mistral (7B)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou 等ACL 2026 · 被引用 51 次
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller 等NDSS 2026 · 被引用 17 次
- InvisibleInk: High-Utility and Low-Cost Text Generation with Differential PrivacyVishnu Vinod, Krishna Pillutla, Abhradeep Guha ThakurtaNeurIPS 2025 · 被引用 12 次
- Extracting alignment data in open modelsFederico Barbero, Xiangming Gu, Christopher A. Choquette Choo, Chawin Sitawarin 等ICML 2026 · 被引用 9 次
- Bypassing Prompt Guards in Production with Controlled-Release PromptingJaiden Fairoze, Sanjam Garg, Keewoo Lee, Mingyuan WangUSENIX Security 2026 · 被引用 8 次
它引用的顶会 Paper14
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach 等ICLR 2022 · 被引用 1,976 次
相关 Paper
- Retracing the Past: LLMs Emit Training Data When They Get LostMyeongseob Ko, Nikhil Reddy Billa, Adam Nguyen, Charles Fleming 等EMNLP 2025 · 被引用 1 次
- Refusal Is Not an Option: Unlearning Safety Alignment of Large Language ModelsMinkyoo Song, Hanna Kim, Jaehan Kim, Seungwon Shin 等USENIX Security 2025
- The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy RisksXiaoyi Chen, Siyuan Tang, Rui Zhu, Shijun Yan 等CCS 2024 · 被引用 18 次
- Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!Zhanhui Zhou, Jie Liu, Zhichen Dong, Jiaheng Liu 等ACL 2024
- Private Memorization Editing: Turning Memorization into a Defense to Strengthen Data Privacy in Large Language ModelsElena Sofia Ruzzetti, Giancarlo A. Xompero, Davide Venditti, Fabio Massimo ZanzottoACL 2025 · 被引用 9 次
