Reconstructing Template-Memorized Images from Natural Prompts
Sol Yarkoni, Mahmood Sharif, Roi Livni
Abstract
Recent advances in generative models, such as diffusion models, have raised several risks and concerns related to privacy, copyright infringement, and data stewardship. To better understand and mitigate these risks, prior work has proposed techniques and attacks that reconstruct images, or parts of images, from the training set. While these approaches demonstrate that training data can be recovered, they often rely on substantial computational resources, access to the training set, or carefully engineered prompts. In this work, we devise a new attack that requires low resources, assumes little to no access to the training data, and identifies seemingly benign prompts that lead to potentially risky image reconstruction. We further show that such reconstructions may occur unintentionally and can be produced by users without specific expertise. For example, we observe that, for one existing model, the prompt "blue Unisex T-Shirt" generates the face of a real individual. Moreover, by combining the identified vulnerabilities with real-world prompt data, we uncover prompts that reproduce memorized elements. Our method builds on intuitions from prior work and leverages domain knowledge to reveal a fundamental vulnerability arising from the use of scraped data from e-commerce platforms, where templated layouts and images are associated with pattern-like prompts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38d1da57-55e9-4b85-a82a-38245bee2b96Builds on15
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Deduplicating Training Data Mitigates Privacy Risks in Language ModelsNikhil Kandpal, Eric Wallace, Colin RaffelICML 2022 · 395 citations
- Understanding and Mitigating Copying in Diffusion ModelsGowthami Somepalli, Vasu Singla, Micah Goldblum, Jonas Geiping et al.NeurIPS 2023 · 265 citations
- Reconstructing Training Data From Trained Neural NetworksNiv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir et al.NeurIPS 2022 · 196 citations
Related papers
- Prompt Stealing Attacks Against Text-to-Image Generation ModelsXinyue Shen, Yiting Qu, Michael Backes, Yang ZhangUSENIX Security 2024 · 65 citations
- Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion ModelsSangwon Jang, June Suk Choi, Jaehyeong Jo, Kimin Lee et al.CVPR 2025
- The Stronger the Diffusion Model, the Easier the Backdoor: Data Poisoning to Induce Copyright BreachesWithout Adjusting Finetuning PipelineHaonan Wang, Qianli Shen, Yao Tong, Yang Zhang et al.ICML 2024 · 48 citations
- SIDE: Surrogate Conditional Data Extraction from Diffusion ModelsYunhao Chen, Shujie Wang, Difan Zou, Xingjun MaAAAI 2026 · 9 citations
- TrojDiff: Trojan Attacks on Diffusion Models with Diverse TargetsWeixin Chen, Dawn Song, Bo LiCVPR 2023
