Lune

NeurIPS2025Top-tier venue

Blameless Users in a Clean Room: Defining Copyright Protection for Generative Models

Aloni Cohen

2025Year
2Citations
1Top-tier citations

Abstract

Are there any conditions under which a generative model's outputs are guaranteed not to infringe the copyrights of its training data? This is the question of "provable copyright protection" first posed in [VKB23]. They define near access-freeness (NAF) and propose it as sufficient for protection. This paper revisits the question and establishes new foundations for provable copyright protection-foundations that are firmer both technically and legally. First, we show that NAF alone does not prevent infringement. In fact, NAF models can enable verbatim copying, a blatant failure of copy protection that we dub being tainted. Then, we introduce our blameless copy protection framework for defining meaningful guarantees, and instantiate it with clean-room copy protection. Clean-room copy protection allows a user to control their risk of copying by behaving in a way that is unlikely to copy in a counterfactual "clean-room setting." Finally, we formalize a common intuition about differential privacy and copyright by proving that DP implies clean-room copy protection when the dataset is golden, a copyright deduplication requirement.

  1. We prove that near access-free (NAF) models can enable verbatim copying. NAF is the first (and only other) attempt at a mathematical definition intended to offer "provable copyright protection" 39th Conference on Neural Information Processing Systems (NeurIPS 2025).

for generative models [VKB23]. The proof leverages NAF's lack of protection against multiple prompts or data-dependent prompts.

  1. We define tainted models, which enable users to reproduce verbatim training data despite knowing nothing about the underlying dataset, a blatant failure of copy protection (Section 5). A meaningful definition of provable copyright protection must exclude tainted models. NAF does not.

  2. We introduce a framework for defining copy protection guarantees, called blameless copy protection (Section 6). It protects blameless users-those who don't themselves induce infringementfrom unwitting copying. 4. We define clean-room copy protection ((κ, β)-clean), a first instantiation of our framework (Section 7), drawing inspiration from clean-room design. A training algorithm is (κ, β)-clean if, for every user who copies in a (counterfactual) "clean-room environment" with probability ≤ β, the probability of copying in the real world is ≤ κ. Clean-room copy protection lets users choose their tolerance for risk κ and then tune β accordingly. Clean-room copy protection also excludes tainted models under mild assumptions.

  3. We prove that differential privacy (DP) implies clean-room copy protection, formalizing a common intuition about DP and copyright (Section 8). This holds when the dataset is golden, a copyright deduplication requirement. Thus, DP provides a way to bring copyrighted expression "into the clean room" without tainting the model. This paper does not offer a legal analysis of where liability would lie under existing law. Instead,our paper develops a mathematical framework for preventing unwitting infringement-where holding the user liable would be unjust-and proposes a rule for assigning fault when copying does inevitably occur (Section 9).

There is a lot of recent work on copyright questions for generative AI [CLG + 23, LCG24, HLJ + 23], and the first generation of cases are working their way through the courts (e.g., New York Times v Microsoft). But most are only tangentially related to goal of stating and proving formal guarantees against copyright infringement.

Vyas, Kakade, and Barak were the first to propose a mathematical property-near access-freeness (NAF)-aimed at "preventing deployment-time copyright infringement" [VKB23]. Elkin-Koren, Hacohen, Livni, and Moran argue that copyright cannot be "reduced to privacy" [EKHLM24]. Referring to both NAF and DP by umbrella term "algorithmic stability", they argue that the sort of provable guarantees that [VKB23] and this paper seek do not capture copyright's complexities. Li, Shen, and Kawaguchi propose and empirically evaluate an attack called VA3 against NAF [LSK24]. They provide good evidence that the CP-k algorithm of [VKB23] may not prevent infringement, but some uncertainty remains (see Appendix B.3). We discuss these three papers at length. There is also work on new NAF algorithms [GAZ + 24, CKOX24]. Scheffler, Tromer, and Varia [STV22] give a complexity-theoretic account of substantial similarity, and a procedure for adjudicating disputes. In contrast, we treat substantial similarity as a black box. That differential privacy might protect against infringement has been suggested in [BLM20, HLJ + 23, VKB23, EKHLM24, CKOX24]. Livni, Moran, Nissim, and Pabbaraju [LMNP24] study the problem of attributing credit to inputs of an algorithm that influence its output, motivated by copyright. They define a condition under which a given input need not be credited (with credit required otherwise). Under their definition, a differentially private

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 4e0293d3-18fa-4226-9af5-99d0fc989406

Cited by top-tier papers1

Ask how each one uses it

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines