DIS-CO: Discovering Copyrighted Content in VLMs Training Data
André V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei Li
Abstract
How can we verify whether copyrighted content was used to train a large vision-language model (VLM) without direct access to its training data? Motivated by the hypothesis that a VLM is able to recognize images from its training corpus, we propose DIS-CO, a novel approach to infer the inclusion of copyrighted content during the model's development. By repeatedly querying a VLM with specific frames from targeted copyrighted material, DIS-CO extracts the content's identity through free-form text completions. To assess its effectiveness, we introduce Movie-Tection, a benchmark comprising 14,000 frames paired with detailed captions, drawn from films released both before and after a model's training cutoff. Our results show that DIS-CO significantly improves detection performance, nearly doubling the average AUC of the best prior method on models with logits available. Our findings also highlight a broader concern: all tested models appear to have been exposed to some extent to copyrighted content. Our code and data are available at https://github.com/avduarte333/ DIS-CO
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 830c4e88-3ae0-4ea7-ab5f-950d5bdbc80fCited by top-tier papers1
Ask how each one uses itBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- ML-Leaks: Model and Data Independent Membership Inference Attacks and Defenses on Machine Learning ModelsAhmed Salem, Yang Zhang, Mathias Humbert, Pascal Berrang et al.NDSS 2019 · 1,141 citations
Related papers
- DE-COP: Detecting Copyrighted Content in Language Models Training DataAndré V. Duarte, Xuandong Zhao, Arlindo L. Oliveira, Lei LiICML 2024 · 81 citations
- Bridging the Copyright Gap: Do Large Vision-Language Models Recognize and Respect Copyrighted Content?Naen Xu, Jinghuai Zhang, Changjiang Li, Hengyu An et al.AAAI 2026 · 6 citations
- Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial ImagesYubo Wang, Jianting Tang, Chaohu Liu, Linli XuICLR 2025
- Detecting Pretraining Data from Large Language ModelsWeijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang et al.ICLR 2024 · 365 citations
- Copyright Traps for Large Language ModelsMatthieu Meeus, Igor Shilov, Manuel Faysse, Yves-Alexandre de MontjoyeICML 2024 · 39 citations
