USENIX Security2025Top-tier venue
On the Proactive Generation of Unsafe Images From Text-To-Image Models Using Benign Prompts
Yixin Wu, Ning Yu, Michael Backes, Yun Shen, Yang Zhang
Abstract
Malicious or manipulated prompts are known to exploit textto-image models to generate unsafe images. Existing studies, however, focus on the passive exploitation of such harmful capabilities. In this paper, we investigate the proactive generation of unsafe images from benign prompts (e.g., a photo of a cat) through maliciously modified text-to-image models. Our preliminary investigation demonstrates that poisoning attacks are a viable method to achieve this goal but uncovers significant side effects, where unintended spread to non-targeted prompts compromises attack stealthiness. Root cause analysis identifies conceptual similarity as an important contributing factor to these side effects. To address this, we propose a stealthy poisoning attack method that balances covertness and performance. Our findings highlight the potential risks of adopting text-to-image models in real-world scenarios, thereby calling for future research and safety measures in this space. 1 Disclaimer. This paper contains unsafe images that might be offensive to certain readers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a3863969-cfc8-4c56-99ef-f2b2f1c8159bCited by top-tier papers10
- Shadowcast: Stealthy Data Poisoning Attacks Against Vision-Language ModelsYuancheng Xu, Jiarui Yao, Manli Shu, Yanchao Sun et al.NeurIPS 2024 · 67 citations
- Safe + Safe = Unsafe? Exploring How Safe Images Can Be Exploited to Jailbreak Large Vision-Language ModelsChenhang Cui, Gelei Deng, An Zhang, Jingnan Zheng et al.NeurIPS 2025 · 10 citations
- SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image ModelsXinfeng Li, Yuchen Yang, Jiangyi Deng, Chen Yan et al.CCS 2024 · 8 citations
- Image-Perfect Imperfections: Safety, Bias, and Authenticity in the Shadow of Text-To-Image Model EvolutionYixin Wu, Yun Shen, Michael Backes, Yang ZhangCCS 2024 · 3 citations
- When Understanding Becomes a Risk: Authenticity and Safety Risks in the Emerging Image Generation ParadigmYe Leng, Junjie Chu, Mingjie Li, Chenhao Lin et al.CVPR 2026 · 3 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- RIATIG: Reliable and Imperceptible Adversarial Text-to-Image Generation with Natural PromptsHan Liu, Yuhao Wu, Shixuan Zhai, Bo Yuan et al.CVPR 2023
- PLA: Prompt Learning Attack Against Text-To-Image Generative ModelsXinqi Lyu, Yihao Liu, Yanjie Li, Bin XiaoICCV 2025 · 10 citations
- USD: NSFW Content Detection for Text-to-Image Models via Scene GraphYuyang Zhang, Kangjie Chen, Xudong Jiang, Jiahui Wen et al.USENIX Security 2025
- Prompt Stealing Attacks Against Text-to-Image Generation ModelsXinyue Shen, Yiting Qu, Michael Backes, Yang ZhangUSENIX Security 2024 · 65 citations
- Towards Human-Imperceptible Backdoor Attacks on Text-to-Image Diffusion ModelsChangkun Wu, Chenghao Chen, Wu kun, Chong Fu et al.CVPR 2026
