Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models
Shawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu, Haitao Zheng, Ben Y. Zhao
Abstract
Trained on billions of images, diffusion-based text-to-image models seem impervious to traditional data poisoning attacks, which typically require poison samples approaching 20% of the training set. In this paper, we show that state-of-the-art text-to-image generative models are in fact highly vulnerable to poisoning attacks. Our work is driven by two key insights. First, while diffusion models are trained on billions of samples, the number of training samples associated with a specific concept or prompt is generally on the order of thousands. This suggests that these models will be vulnerable to prompt-specific poisoning attacks that corrupt a model’s ability to respond to specific targeted prompts. Second, poison samples can be carefully crafted to maximize poison potency to ensure success with very few samples.We introduce Nightshade, a prompt-specific poisoning attack optimized for potency that can completely control the output of a prompt in Stable Diffusion’s newest model (SDXL) with less than 100 poisoned training samples. Nightshade also generates stealthy poison images that look visually identical to their benign counterparts, and produces poison effects that "bleed through" to related concepts. More importantly, a moderate number of Nightshade attacks on independent prompts can destabilize a model and disable its ability to generate images for any and all prompts. Finally, we propose the use of Nightshade and similar tools as a defense for content owners against web scrapers that ignore opt-out/do-not-crawl directives, and discuss potential implications for both model trainers and content owners.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 12a7d1e0-fbeb-42f9-8a67-e9558a168ddbCited by top-tier papers34
- Governance of Generative AI in Creative Work: Consent, Credit, Compensation, and BeyondLin Kyi, Amruta Mahuli, Michael Six Silberman, Reuben Binns et al.CHI 2025 · 50 citations
- Generative AI Uses and Risks for Knowledge Workers in a Science OrganizationKelly B. Wagman, Matthew T. Dearing, Marshini ChettyCHI 2025 · 16 citations
- Creativity Supportive Ecosystems: A Framework for Understanding Function and Disruption in Online Art WorldsShm Garanganao Almeda, Joy O. Kim, Bjoern HartmannCHI 2025 · 13 citations
- Copying style, Extracting value: Illustrators' Perception of AI Style Transfer and its Impact on Creative LaborJulien Porquet, Sitong Wang, Lydia B. ChiltonCHI 2025 · 12 citations
- Disguised Copyright Infringement of Latent Diffusion ModelsYiwei Lu, Matthew Y. R. Yang, Zuoqiu Liu, Gautam Kamath et al.ICML 2024 · 10 citations
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- LightShed: Defeating Perturbation-based Image Copyright ProtectionsHanna Foerster, Sasha Behrouzi, Phillip Rieger, Murtuza Jadliwala et al.USENIX Security 2025
- Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion ModelsSangwon Jang, June Suk Choi, Jaehyeong Jo, Kimin Lee et al.CVPR 2025
- Understanding Implosion in Text-to-Image Generative ModelsWenxin Ding, Cathy Yuanchen Li, Shawn Shan, Ben Y. Zhao et al.CCS 2024 · 2 citations
- Latent Diffusion Unlearning: Protecting Against Unauthorized Personalization Through Trajectory Shifted PerturbationsNaresh Kumar Devulapally, Shruti Agarwal, Tejas Gokhale, Vishnu Suresh LokhandeACM MM 2025 · 1 citation
- SafeGuider: Robust and Practical Content Safety Control for Text-to-Image ModelsPeigui Qi, Kunsheng Tang, Wenbo Zhou, Weiming Zhang et al.CCS 2025
