Copyright Traps for Large Language Models
Matthieu Meeus, Igor Shilov, Manuel Faysse, Yves-Alexandre de Montjoye
Abstract
Questions of fair use of copyright-protected content to train Large Language Models (LLMs) are being actively debated. Document-level inference has been proposed as a new task: inferring from black-box access to the trained model whether a piece of content has been seen during training. SOTA methods however rely on naturally occurring memorization of (part of) the content. While very effective against models that memorize significantly, we hypothesize--and later confirm--that they will not work against models that do not naturally memorize, e.g. medium-size 1B models. We here propose to use copyright traps, the inclusion of fictitious entries in original content, to detect the use of copyrighted materials in LLMs with a focus on models where memorization does not naturally occur. We carefully design a randomized controlled experimental setup, inserting traps into original content (books) and train a 1.3B LLM from scratch. We first validate that the use of content in our target model would be undetectable using existing methods. We then show, contrary to intuition, that even medium-length trap sentences repeated a significant number of times (100) are not detectable using existing methods. However, we show that longer sequences repeated a large number of times can be reliably detected (AUC=0.75) and used as copyright traps. Beyond copyright applications, our findings contribute to the study of LLM memorization: the randomized controlled setup enables us to draw causal relationships between memorization and certain sequence properties such as repetition in model training data and perplexity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9461b008-6360-4255-97fa-3a8e2e7844f8Cited by top-tier papers14
- Exploring the limits of strong membership inference attacks on large language modelsJamie Hayes, Ilia Shumailov, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2025 · 26 citations
- Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data PoisoningWassim Bouaziz, Mathurin Videau, Nicolas Usunier, El-Mahdi El-MhamdiICLR 2026 · 8 citations
- Free and Fair Hardware: A Pathway to Copyright Infringement-Free Verilog Generation using LLMsSam Bush, Matthew DeLorenzo, Phat Tieu, Jeyavijayan RajendranDAC 2025 · 5 citations
- Dual-Space Smoothness for Robust and Balanced LLM UnlearningHan Yan, Zheyuan Liu, Meng JiangICLR 2026 · 5 citations
- Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory ProbingJinhua Yin, Peiru Yang, Chen Yang, Huili Wang et al.NeurIPS 2025 · 4 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
Related papers
- DE-COP: Detecting Copyrighted Content in Language Models Training DataAndré V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei LiICML 2024 · 81 citations
- LLM Dataset Inference: Did you train on my dataset?Pratyush Maini, Hengrui Jia, Nicolas Papernot, Adam DziedzicNeurIPS 2024 · 162 citations
- Data Provenance Auditing of Fine-Tuned Large Language Models with a Text-Preserving TechniqueYanming Li, Cédric Eichler, Nicolas Anciaux, Alexandra Bensamoun et al.ICML 2026 · 3 citations
- Did the Neurons Read your Book? Document-level Membership Inference for Large Language ModelsMatthieu Meeus, Shubham Jain, Marek Rei, Yves-Alexandre de MontjoyeUSENIX Security 2024 · 67 citations
- Perturb Your Data: Paraphrase-Guided Training Data WatermarkingPranav Shetty, Mirazul Haque, Petr Babkin, Zhiqiang Ma et al.AAAI 2026
