Diffusion Art or Digital Forgery? Investigating Data Replication in Diffusion Models
Gowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping, Tom Goldstein
Abstract
Figure 1. Stable Diffusion is capable of reproducing training data, creating images by piecing together foreground and background objects that it has memorized. Furthermore, the system sometimes exhibits reconstructive memory, in which recalled objects are semantically equivalent to their source object without being pixel-wise identical. Here, we show this behavior occurring with a range of prompts sampled from LAION, and with a hand-crafted prompt (rightmost pair). The presence of such images raises questions about the nature of data memorization and the ownership of diffusion images. Top row: generated images. Bottom row: closest matches in the LAION-Aesthetics v2 6+ set. Sometimes source and match prompts are quite similar, and sometimes they are quite different. See Figure 7 for more examples with prompts, or the Appendix for prompts from this figure.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02ee8b2a-c9fa-4c26-8ced-2b499ffc6f74Cited by top-tier papers173
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and GenerationChongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong et al.ICLR 2024 · 351 citations
- Ablating Concepts in Text-to-Image Diffusion ModelsNupur Kumari, Bingliang Zhang, Sheng-Yu Wang, Eli Shechtman et al.ICCV 2023 · 327 citations
- Understanding and Mitigating Copying in Diffusion ModelsGowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping et al.NeurIPS 2023 · 265 citations
- Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion modelsGeorge Stein, Jesse C. Cresswell, Rasa Hosseinzadeh, Yi Sui et al.NeurIPS 2023 · 260 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
Related papers
- Extracting Training Data from Diffusion ModelsNicholas Carlini, Jamie Hayes, Milad Nasr, Matthew Jagielski et al.USENIX Security 2023
- Broken Memories: Detecting and Mitigating Memorization in Diffusion Models with Degraded GenerationsYuanmin Huang, Mi Zhang, Chen Chen, Feifei Li et al.KDD 2026
- The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs BetterScott Geng, Cheng-Yu Hsieh, Vivek Ramanujan, Matthew Wallingford et al.NeurIPS 2024 · 27 citations
- Enhancing Compositional Text-to-Image Generation with Reliable Random SeedsShuangqi Li, Hieu Le, Jingyi Xu, Mathieu SalzmannICLR 2025
- You Don’t Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion ModelsKairan Zhao, Eleni Triantafillou, Peter TriantafillouICML 2026
