Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?
Yuechen Xie, Jie Song, Huiqiong Wang, Mingli Song
摘要
High-quality open-source text-to-image models have lowered the threshold for obtaining photorealistic images significantly, but also face potential risks of misuse. Specifically, suspects may use synthetic data generated by these generative models to train models for specific tasks without permission, when lacking real data resources especially. Protecting these generative models is crucial for the wellbeing of their owners. In this work, we propose the first method to this important yet unresolved issue, called Training data Provenance Verification (TrainProVe). The rationale behind TrainProVe is grounded in the principle of generalization error bound, which suggests that, for two models with the same task, if the distance between their training data distributions is smaller, their generalization ability will be closer. We validate the efficacy of Train-ProVe across four text-to-image models (Stable Diffusion v1.4, latent consistency model, and Stable Cascade). The results show that TrainProVe achieves a verification accuracy of over 99% in determining the provenance of suspicious model training data, surpassing all previous methods. Code is available at https://github.com/ xieyc99/TrainProVe .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Exploring the Underwater World Segmentation without Extra TrainingBingyu Li, Tao Huo, Da Zhang, Zhiyuan Zhao 等CVPR 2026 · 被引用 18 次
- Dataset Ownership Verification for Pre-Trained Masked ModelsYuechen Xie, Jie Song, Yicheng Shan, Xiaoyan Zhang 等ICCV 2025 · 被引用 1 次
- MARIS: Marine Open-Vocabulary Instance SegmentationBingyu Li, Feiyu Wang, Da Zhang, Zhiyuan Zhao 等CVPR 2026
- Dataset Ownership Verification in Contrastive Pre-trained ModelsYuechen Xie, Jie Song, Mengqi Xue, Haofei Zhang 等ICLR 2025
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
相关 Paper
- Training Data Attribution: Was Your Model Secretly Trained On Data Created By Mine?Likun Zhang, Hao Wu, Lingcui Zhang, Fengyuan Xu 等KDD 2025
- How to Trace Latent Generative Model Generated Images without Artificial Watermark?Zhenting Wang, Vikash Sehwag, Chen Chen, Lingjuan Lyu 等ICML 2024 · 被引用 24 次
- DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation ModelsZeyang Sha, Zheng Li, Ning Yu, Yang ZhangCCS 2023 · 被引用 123 次
- FakeInversion: Learning to Detect Images from Unseen Text-to-Image Models by Inverting Stable DiffusionGeorge Cazenavette, Avneesh Sud, Thomas Leung, Ben UsmanCVPR 2024
- Ghost in the Transformer: Detecting Model Reuse with Invariant Spectral SignaturesSuqing Wang, Ziyang Ma, Xinyi Li, Zuchao LiAAAI 2026 · 被引用 1 次
