On Provable Copyright Protection for Generative Models
Nikhil Vyas, Sham M. Kakade, Boaz Barak
Abstract
There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data that was in their training set. We give a formal definition of and prove bounds on the probability that a model satisfying this definition outputs a sample similar to , even if is included in its training set. Roughly speaking, a generative model is \textit{k-NAF} if for every potentially copyrighted data , the output of diverges by at most -bits from the output of a model that \textit{did not access C at all}. We also give generative model learning algorithms, which efficiently modify the original generative model learning algorithm in a black box manner, that output generative models with strong bounds on the probability of sampling protected content. Furthermore, we provide promising experiments for both language (transformers) and image (diffusion) generative models, showing minimal degradation in output quality while ensuring strong protections against sampling protected content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d0b457fd-9b5d-48c8-8d7a-709ff1079748Cited by top-tier papers34
- Understanding and Mitigating Copying in Diffusion ModelsGowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping et al.NeurIPS 2023 · 265 citations
- On the Generalization Properties of Diffusion ModelsPuheng Li, Zhong Li, Huishuai Zhang, Jiang BianNeurIPS 2023 · 86 citations
- DE-COP: Detecting Copyrighted Content in Language Models Training DataAndré V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei LiICML 2024 · 81 citations
- On the Edge of Memorization in Diffusion ModelsSam Buchanan, Druv Pai, Yi Ma, Valentin De BortoliNeurIPS 2025 · 25 citations
- Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data?Michael-Andrei Panaitescu-Liess, Zora Che, Bang An, Yuancheng Xu et al.AAAI 2025 · 21 citations
Builds on11
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Large Language Models Can Be Strong Differentially Private LearnersXuechen Li, Florian Tramèr, Percy Liang, Tatsunori HashimotoICLR 2022 · 502 citations
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
- Differentially Private Learning Needs Better Features (or Much More Data)Florian Tramèr, Dan BonehICLR 2021 · 325 citations
Related papers
- Blameless Users in a Clean Room: Defining Copyright Protection for Generative ModelsAloni CohenNeurIPS 2025 · 2 citations
- CPR: Retrieval Augmented Generation for Copyright ProtectionAditya Golatkar, Alessandro Achille, Luca Zancato, Yu-Xiang Wang et al.CVPR 2024
- Disguised Copyright Infringement of Latent Diffusion ModelsYiwei Lu, Matthew Y. R. Yang, Zuoqiu Liu, Gautam Kamath et al.ICML 2024 · 10 citations
- Watermark-embedded Adversarial Examples for Copyright Protection against Diffusion ModelsPeifei Zhu, Tsubasa Takahashi, Hirokatsu KataokaCVPR 2024 · 14 citations
- Training-Free Safe Denoisers for Safe Use of Diffusion ModelsMingyu Kim, Dongjun Kim, Amman Yusuf, Stefano Ermon et al.NeurIPS 2025 · 21 citations
