Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4
Kent K. Chang, Mackenzie Cramer, Sandeep Soni, David Bamman
摘要
In this work, we carry out a data archaeology to infer books that are known to ChatGPT and GPT-4 using a name cloze membership inference query. We find that OpenAI models have memorized a wide collection of copyrighted materials, and that the degree of memorization is tied to the frequency with which passages of those books appear on the web. The ability of these models to memorize an unknown set of books complicates assessments of measurement validity for cultural analytics by contaminating test data; we show that models perform much better on memorized books than on non-memorized books for downstream tasks. We argue that this supports a case for open models whose training data is known.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper40
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 被引用 173 次
- No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model PerformanceVishaal Udandarao, Ameya Prabhu, Adhiraj Ghosh, Yash Sharma 等NeurIPS 2024 · 被引用 101 次
- On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMsJen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam 等ICLR 2024 · 被引用 85 次
- Apertus: Democratizing Open and Compliant LLMs for Global Language EnvironmentsAlejandro Hernández-Cano, Alexander Hägele, Allen Hao Huang, Angelika Romanou 等ACL 2026 · 被引用 51 次
- To the Cutoff... and Beyond? A Longitudinal Perspective on LLM Data ContaminationManley Roberts, Himanshu Thakur, Christine Herlihy, Colin White 等ICLR 2024 · 被引用 45 次
它引用的顶会 Paper13
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 被引用 5,137 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Generalization through Memorization: Nearest Neighbor Language ModelsUrvashi Khandelwal, Omer Levy, Dan Jurafsky, Luke Zettlemoyer 等ICLR 2020 · 被引用 1,038 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 被引用 395 次
相关 Paper
- Copyright Traps for Large Language ModelsMatthieu Meeus, Igor Shilov, Manuel Faysse, Yves-Alexandre de MontjoyeICML 2024 · 被引用 39 次
- LLM Dataset Inference: Did you train on my dataset?Pratyush Maini, Hengrui Jia, Nicolas Papernot, Adam DziedzicNeurIPS 2024 · 被引用 162 次
- Did the Neurons Read your Book? Document-level Membership Inference for Large Language ModelsMatthieu Meeus, Shubham Jain, Marek Rei, Yves-Alexandre de MontjoyeUSENIX Security 2024 · 被引用 67 次
- Counterfactual Memorization in Neural Language ModelsChiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski 等NeurIPS 2023 · 被引用 184 次
- Scalable Extraction of Training Data from Aligned, Production Language ModelsMilad Nasr, Javier Rando, Nicholas Carlini, Jonathan Hayase 等ICLR 2025
