Lune

ICML2023Top-tier venue

On Provable Copyright Protection for Generative Models

Nikhil Vyas, Sham M. Kakade, Boaz Barak

2023Year
120Citations
34Top-tier citations

Abstract

There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data CC that was in their training set. We give a formal definition of near access-freeness (NAF)\textit{near access-freeness (NAF)} and prove bounds on the probability that a model satisfying this definition outputs a sample similar to CC, even if CC is included in its training set. Roughly speaking, a generative model pp is \textit{k-NAF} if for every potentially copyrighted data CC, the output of pp diverges by at most kk-bits from the output of a model qq that \textit{did not access C at all}. We also give generative model learning algorithms, which efficiently modify the original generative model learning algorithm in a black box manner, that output generative models with strong bounds on the probability of sampling protected content. Furthermore, we provide promising experiments for both language (transformers) and image (diffusion) generative models, showing minimal degradation in output quality while ensuring strong protections against sampling protected content.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d0b457fd-9b5d-48c8-8d7a-709ff1079748

Cited by top-tier papers34

Ask how each one uses it

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines