USENIX Security2026Top-tier venue
Identifying Provenance of Generative Text-to-Image Models
Anna Yoo Jeong Ha, Wenxin Ding, Stanley Wu, Shawn Shan, Haitao Zheng, Ben Y. Zhao
Abstract
Fine-tuning provides a fast and cheap way to produce new text-to-image models that are often indistinguishable from ones trained from scratch. Unfortunately, misrepresentation of fine-tuned models creates problems for AI companies and users alike, by disincentivizing competition and misleading users on model quality and ethics of its training process.
In this paper, we propose a model provenance system that identifies models produced by fine-tuning on existing base text-to-image models, using only black-box query access to the models. Our design is informed by analysis showing that one can quantify the feature space difference between textto-image models by analyzing their responses to detailed prompts. Given a target model, our system analyzes its output, extracts visual features using a generic feature extractor, and compares the distribution against those derived from a pool of base models using Jensen-Shannon divergence. We then apply statistical hypothesis testing to determine if the target model is trained from scratch or fine-tuned, and if the latter, the likely base (parent) model. We evaluate our system across seven popular diffusion models and numerous fine-tuned variants. Our results show high accuracy in attributing model lineage, even under adversarial conditions such as image postprocessing or weight perturbations. Finally, we demonstrate real-world efficacy of our system by tracing provenance of in-the-wild models from popular online platforms.
Model trainer fine-tunes and claims ownership of the fine-tuned model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 99e235cb-d8bf-4d97-97ba-2225ada37e72Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Model Provenance Testing for Large Language ModelsIvica Nikolic, Teodora Baluta, Prateek SaxenaNeurIPS 2025 · 20 citations
- Matching Pairs: Attributing Fine-Tuned Models to their Pre-Trained Large Language ModelsMyles Foley, Ambrish Rawat, Taesung Lee, Yufang Hou et al.ACL 2023 · 2 citations
- Training Data Provenance Verification: Did Your Model Use Synthetic Data from My Generative Model for Training?Yuechen Xie, Jie Song, Huiqiong Wang, Mingli SongCVPR 2025
- DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation ModelsZeyang Sha, Zheng Li, Ning Yu, Yang ZhangCCS 2023 · 123 citations
- CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image ModelsJunhoo Lee, Mijin Koo, Nojun KwakCVPR 2026
