Training Data Attribution: Was Your Model Secretly Trained On Data Created By Mine?
Likun Zhang, Hao Wu, Lingcui Zhang, Fengyuan Xu, Jin Cao, Fenghua Li, Ben Niu
Abstract
The emergence of text-to-image models has recently sparked significant interest, but the attendant is a looming shadow of potential infringement by violating user terms. Specifically, an adversary may exploit data created by a commercial model to train their own without proper authorization. To address such risk, it is crucial to investigate the attribution of a suspicious model's training data by determining whether its training data originates, wholly or partially, from a specific source model. To trace the generated data, existing methods need to apply additional watermarks during either the training or inference phases of the source model. However, these methods are impractical for pre-trained models that have been released, especially when model owners lack security expertise. To tackle this challenge, we propose an injection-free training data attribution method for text-to-image models. It can identify whether a model's training data stems from a certain source model without adding additional watermarks on the source model. The rationale of our method lies in the inherent memorization characteristic of text-to-image models. The memorization of training data is inherited through the data generated by the source model to the model trained on that data, making the source model and the infringing model exhibit consistent behaviors on specific samples. Therefore, from instance-level, we develop detection-based and generation-based strategies to uncover these distinct samples and using them as inherent watermarks to verify if a suspicious model originates from the source model. Besides, we also propose a statistical-level attribution method, utilizing the shadow model technique to train an attribution discriminator. Experiments demonstrate that the attribution accuracy and AUC scores of our methods are over 80% even when the infringing model only uses a small proportion of generated data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 416d571c-87e0-4038-a7d4-cf93cd553f70Cited by top-tier papers2
- Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?Yuechen Xie, Jie Song, Huiqiong Wang, Mingli SongCVPR 2025
- On the Fragility of Data Attribution When Learning Is DistributedXian Gao, Bo Hui, MIN-TE SUN, Wei-Shinn KuICML 2026
Builds on14
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- Membership Inference Attacks From First PrinciplesNicholas Carlini, Steve Chien, Milad Nasr, Shuang Song et al.S&P 2022 · 1,049 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Entangled Watermarks as a Defense against Model ExtractionHengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, Nicolas PapernotUSENIX Security 2021 · 287 citations
Related papers
- DIAGNOSIS: Detecting Unauthorized Data Usages in Text-to-image Diffusion ModelsZhenting Wang, Chen Chen, Lingjuan Lyu, Dimitris N. Metaxas et al.ICLR 2024 · 12 citations
- Disentangled Style Domain for Implicit z-Watermark Towards Copyright ProtectionJunqiang Huang, Zhaojun Guo, Ge Luo, Zhenxing Qian et al.NeurIPS 2024 · 5 citations
- Where Did I Come From? Origin Attribution of AI-Generated ImagesZhenting Wang, Chen Chen, Yi Zeng, Lingjuan Lyu et al.NeurIPS 2023 · 44 citations
- Image-level Memorization Detection via Inversion-based Inference PerturbationYue Jiang, Haokun Lin, Yang Bai, Bo Peng et al.ICLR 2025
- ArtistAuditor: Auditing Artist Style Pirate in Text-to-Image Generation ModelsLinkang Du, Zheng Zhu, Min Chen, Zhou Su et al.WWW 2025 · 5 citations
