Bag of Tricks for Training Data Extraction from Language Models
Weichen Yu, Tianyu Pang, Qian Liu, Chao Du, Bingyi Kang, Yan Huang, Min Lin, Shuicheng Yan
Abstract
With the advance of language models, privacy protection is receiving more attention. Training data extraction is therefore of great importance, as it can serve as a potential tool to assess privacy leakage. However, due to the difficulty of this task, most of the existing methods are proof-of-concept and still not effective enough. In this paper, we investigate and benchmark tricks for improving training data extraction using a publicly available dataset. Because most existing extraction methods use a pipeline of generating-then-ranking, i.e., generating text candidates as potential training data and then ranking them based on specific criteria, our research focuses on the tricks for both text generation (e.g., sampling strategy) and text ranking (e.g., token-level criteria). The experimental results show that several previously overlooked tricks can be crucial to the success of training data extraction. Based on the GPT-Neo 1.3B evaluation results, our proposed tricks outperform the baseline by a large margin in most cases, providing a much stronger baseline for future research. The code is available at https://github.com/weichen-yu/LM-Extraction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2eea9ba6-7d65-450a-bb45-71a98b4dcb8dCited by top-tier papers18
- DNA-GPT: Divergent N-Gram Analysis for Training-Free Detection of GPT-Generated TextXianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold et al.ICLR 2024 · 173 citations
- LLM-PBE: Assessing Data Privacy in Large Language ModelsQinbin Li, Junyuan Hong, Chulin Xie, Jeffrey Tan et al.VLDB 2024 · 66 citations
- Large Language Models Can Be Contextual Privacy Protection LearnersYijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu et al.EMNLP 2024 · 18 citations
- MoPe: Model Perturbation based Privacy Attacks on Language ModelsMarvin Li, Jason Wang, Jeffrey G. Wang, Seth NeelEMNLP 2023 · 8 citations
- Promise and Peril of Collaborative Code Generation Models: Balancing Effectiveness and MemorizationZhi Chen, Lingxiao JiangASE 2024 · 4 citations
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Stealing Machine Learning Models via Prediction APIsFlorian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter et al.USENIX Security 2016 · 2,088 citations
Related papers
- Private Investigator: Extracting Personally Identifiable Information from Large Language Models Using Optimized PromptsSeongho Keum, Dongwon Shin, Leo Marchyok, Sanghyun Hong et al.USENIX Security 2025
- Scalable Extraction of Training Data from Aligned, Production Language ModelsMilad Nasr, Javier Rando, Nicholas Carlini, Jonathan Hayase et al.ICLR 2025
- Investigating How Pre-training Data Leakage Affects Models' Reproduction and Detection CapabilitiesMasahiro Kaneko, Timothy BaldwinEMNLP 2025 · 2 citations
- Codebreaker: Dynamic Extraction Attacks on Code Language ModelsChangzhou Han, Zehang Deng, Wanlun Ma, Xiaogang Zhu et al.S&P 2025
- Effective PII Extraction from LLMs through Augmented Few-Shot LearningShuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang et al.USENIX Security 2025
