Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image Models
Yiting Qu, Xinyue Shen, Xinlei He, Michael Backes, Savvas Zannettou, Yang Zhang
Abstract
State-of-the-art Text-to-Image models like Stable Diffusion and DALLE2 are revolutionizing how people generate visual content. At the same time, society has serious concerns about how adversaries can exploit such models to generate problematic or unsafe images. In this work, we focus on demystifying the generation of unsafe images and hateful memes from Text-to-Image models. We first construct a typology of unsafe images consisting of five categories (sexually explicit, violent, disturbing, hateful, and political). Then, we assess the proportion of unsafe images generated by four advanced Text-to-Image models using four prompt datasets. We find that Text-to-Image models can generate a substantial percentage of unsafe images; across four models and four prompt datasets, 14.56% of all generated images are unsafe. When comparing the four Text-to-Image models, we find different risk levels, with Stable Diffusion being the most prone to generating unsafe content (18.92% of all generated images are unsafe). Given Stable Diffusion's tendency to generate more unsafe content, we evaluate its potential to generate hateful meme variants if exploited by an adversary to attack a specific individual or community. We employ three image editing methods, DreamBooth, Textual Inversion, and SDEdit, which are supported by Stable Diffusion to generate variants. Our evaluation result shows that 24% of the generated images using DreamBooth are hateful meme variants that present the features of the original hateful meme and the target individual/community; these generated images are comparable to hateful meme variants collected from the real world. Overall, our results demonstrate that the danger of large-scale generation of unsafe images is imminent. We discuss several mitigating measures, such as curating training data, regulating prompts, and implementing safety filters, and encourage better safeguard tools to be developed to prevent unsafe generation.1 Our code is available at https://github.com/YitingQu/unsafe-diffusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb5d8440-1ca2-4cdd-9e29-9ddf7a34221bCited by top-tier papers80
- Ring-A-Bell! How Reliable are Concept Removal Methods For Diffusion Models?Yu-Lin Tsai, Chia-Yi Hsu, Chulin Xie, Chih-Hsun Lin et al.ICLR 2024 · 207 citations
- SneakyPrompt: Jailbreaking Text-to-image Generative ModelsYuchen Yang, Bo Hui, Haolin Yuan, Neil Gong et al.S&P 2024 · 188 citations
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen et al.CCS 2024 · 132 citations
- Circumventing Concept Erasure Methods For Text-To-Image Generative ModelsMinh Pham, Kelly O. Marshall, Niv Cohen, Govind Mittal et al.ICLR 2024 · 82 citations
- GuardT2I: Defending Text-to-Image Models from Adversarial PromptsYijun Yang, Ruiyuan Gao, Xiao Yang, Jianyuan Zhong et al.NeurIPS 2024 · 74 citations
Builds on13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- SDEdit: Guided Image Synthesis and Editing with Stochastic Differential EquationsChenlin Meng, Yutong He, Yang Song, Jiaming Song et al.ICLR 2022 · 2,128 citations
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang et al.ICLR 2022 · 753 citations
Related papers
- SurrogatePrompt: Bypassing the Safety Filter of Text-to-Image Models via SubstitutionZhongjie Ba, Jieming Zhong, Jiachen Lei, Peng Cheng et al.CCS 2024 · 7 citations
- Multimodal Pragmatic Jailbreak on Text-to-image ModelsTong Liu, Zhixin Lai, Jiawen Wang, Gengyuan Zhang et al.ACL 2025
- SafeGen: Mitigating Sexually Explicit Content Generation in Text-to-Image ModelsXinfeng Li, Yuchen Yang, Jiangyi Deng, Chen Yan et al.CCS 2024 · 8 citations
- DE-FAKE: Detection and Attribution of Fake Images Generated by Text-to-Image Generation ModelsZeyang Sha, Zheng Li, Ning Yu, Yang ZhangCCS 2023 · 123 citations
- ZeroFake: Zero-Shot Detection of Fake Images Generated and Edited by Text-to-Image Generation ModelsZeyang Sha, Yicong Tan, Mingjie Li, Michael Backes et al.CCS 2024 · 8 citations
