Certified Mitigation of Worst-Case LLM Copyright Infringement
Jingyu Zhang, Jiacan Yu, Marc Marone, Benjamin Van Durme, Daniel Khashabi
摘要
The exposure of large language models (LLMs) to copyrighted material during pre-training raises concerns about unintentional copyright infringement post deployment. This has driven the development of "copyright takedown" methods-post-training approaches aimed at preventing models from generating content substantially similar to copyrighted ones. While current mitigation approaches are somewhat effective for average-case risks, we demonstrate that they overlook worst-case copyright risks exhibited by the existence of long, verbatim quotes from copyrighted sources. We propose BLOOMSCRUB, a remarkably simple yet highly effective inference-time approach that provides certified copyright takedown. Our method repeatedly interleaves quote detection with rewriting techniques to transform potentially infringing segments. By leveraging efficient data sketches (Bloom filters), our approach enables scalable copyright screeningeven for large-scale real-world corpora. When quotes beyond a length threshold cannot be removed, the system can abstain from responding, offering certified risk reduction. Experimental results show that BLOOMSCRUB reduces infringement risk, preserves utility, and accommodates different levels of enforcement stringency with adaptive abstention. Our results suggest that lightweight, inference-time methods can be surprisingly effective for copyright prevention. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SCOPE: Intrinsic Semantic Space Control for Mitigating Copyright Infringement in LLMsZhenliang Zhang, Xinyu Hu, Xiaojun WanAAAI 2026 · 被引用 1 次
- Anchored Decoding: Provably Reducing Copyright Risk for Any Language ModelJacqueline He, Jonathan Hayase, Scott Yih, Sewoong Oh 等ICML 2026
它引用的顶会 Paper6
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Emergent and Predictable Memorization in Large Language ModelsStella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf 等NeurIPS 2023 · 被引用 205 次
- Quantifying Memorization Across Neural Language ModelsNicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee 等ICLR 2023 · 被引用 158 次
- Speak, Memory: An Archaeology of Books Known to ChatGPT/GPT-4Kent K. Chang, Mackenzie Cramer, Sandeep Soni, David BammanEMNLP 2023 · 被引用 70 次
- Be like a Goldfish, Don't Memorize! Mitigating Memorization in Generative LLMsAbhimanyu Hans, John Kirchenbauer, Yuxin Wen, Neel Jain 等NeurIPS 2024 · 被引用 65 次
相关 Paper
- Large Language Model Unlearning for Source CodeXue Jiang, Yihong Dong, Huangzhao Zhang, Tangxinyu Wang 等AAAI 2026
- Copyright-Protected Language Generation via Adaptive Model FusionJavier Abad, Konstantin Donhauser, Francesco Pinto, Fanny YangICLR 2025
- DE-COP: Detecting Copyrighted Content in Language Models Training DataAndré V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei LiICML 2024 · 被引用 81 次
- SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text GenerationXiaoze Liu, Ting Sun, Tianyang Xu, Feijie Wu 等EMNLP 2024 · 被引用 4 次
- Copyright Traps for Large Language ModelsMatthieu Meeus, Igor Shilov, Manuel Faysse, Yves-Alexandre de MontjoyeICML 2024 · 被引用 39 次
