Measuring Non-Adversarial Reproduction of Training Data in Large Language Models
Michael Aerni, Javier Rando, Edoardo Debenedetti, Nicholas Carlini, Daphne Ippolito, Florian Tramèr
摘要
Large language models memorize parts of their training data. Memorizing short snippets and facts is required to answer questions about the world and to be fluent in any language. But models have also been shown to reproduce long verbatim sequences of memorized text when prompted by a motivated adversary. In this work, we investigate an intermediate regime of memorization that we call non-adversarial reproduction, where we quantify the overlap between model responses and pretraining data when responding to natural and benign prompts. For a variety of innocuous prompt categories (e.g., writing a letter or a tutorial), we show that up to 15% of the text output by popular conversational language models overlaps with snippets from the Internet. In worst cases, we find generations where 100% of the content can be found exactly online. For the same tasks, we find that human-written text has far less overlap with Internet data. We further study whether prompting strategies can close this reproduction gap between models and humans. While appropriate prompting can reduce non-adversarial reproduction on average, we find that mitigating worst-case reproduction of training data requires stronger defenses -- even for benign interactions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Blackbox Model Provenance via Palimpsestic Membership InferenceRohith Kuditipudi, Jing Huang, Sally Zhu, Diyi Yang 等NeurIPS 2025 · 被引用 12 次
- Measuring LLM Novelty As The Frontier Of Original And High-Quality OutputVishakh Padmakumar, Chen Yueh-Han, Jane Pan, Valerie Chen 等ICLR 2026 · 被引用 6 次
- TokenSwap: A Lightweight Method to Disrupt Memorized Sequences in LLMsParjanya Prajakta Prashant, Kaustubh Ponkshe, Babak SalimiNeurIPS 2025 · 被引用 2 次
- Demystifying Foreground-Background Memorization in Diffusion ModelsJimmy Z. Di, Yiwei Lu, Yaoliang Yu, Gautam Kamath 等AAAI 2026 · 被引用 1 次
- DefGen-Bench: A Benchmark for Chinese Criminal Defence Opinion Generation in LegalAISenbo Zhang, Qiqi Wang, Fanghao Lou, Guanyu Chen 等ACL 2026
它引用的顶会 Paper13
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
相关 Paper
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- Rethinking LLM Memorization through the Lens of Adversarial CompressionAvi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C. Lipton 等NeurIPS 2024 · 被引用 120 次
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 被引用 395 次
- Quantifying and Analyzing Entity-Level Memorization in Large Language ModelsZhenhong Zhou, Jiuyang Xiang, Chaomeng Chen, Sen SuAAAI 2024 · 被引用 23 次
- Counterfactual Memorization in Neural Language ModelsChiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski 等NeurIPS 2023 · 被引用 184 次
