Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
Yiting Qu, Ziqing Yang, Yihan Ma, Michael Backes, Savvas Zannettou, Yang Zhang
Abstract
Recent advances in text-to-image diffusion models have enabled the creation of a new form of digital art: optical illusions—visual tricks that create different perceptions of reality. However, adversaries may misuse such techniques to generate hateful illusions, which embed specific hate messages into harmless scenes and disseminate them across web communities. In this work, we take the first step toward investigating the risks of scalable hateful illusion generation and the potential for bypassing current content moderation models. Specifically, we generate 1,860 optical illusions using Stable Diffusion and ControlNet, conditioned on 62 hate messages. Of these, 1,571 are hateful illusions that successfully embed hate messages, either overtly or subtly, forming the Hateful Illusion dataset. Using this dataset, we evaluate the performance of six moderation classifiers and nine vision language models (VLMs) in identifying hateful illusions. Experimental results reveal significant vulnerabilities in existing moderation models: the detection accuracy falls below 0.245 for moderation classifiers and below 0.102 for VLMs. We further identify a critical limitation in their vision encoders, which mainly focus on surface-level image details while overlooking the secondary layer of information, i.e., hidden messages. To address this risk, we explore preliminary mitigation measures and identify the most effective approaches from the perspectives of image transformations and training-level strategies. 11Our code is available at https://github.com/TrustAIRLab/HatefulIllusion. Disclaimer. This paper contains hateful images. Reader discretion is advised.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aae640e4-36f1-4fe0-a2a1-bafe2e55e800Cited by top-tier papers1
Ask how each one uses itBuilds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
Related papers
- From Meme to Threat: On the Hateful Meme Understanding and Induced Hateful Content Generation in Open-Source Vision Language ModelsYihan Ma, Xinyue Shen, Yiting Qu, Ning Yu et al.USENIX Security 2025
- Unsafe Diffusion: On the Generation of Unsafe Images and Hateful Memes From Text-To-Image ModelsYiting Qu, Xinyue Shen, Xinlei He, Michael Backes et al.CCS 2023 · 48 citations
- What the Eyes See, the LLMs Miss: Exploiting Human Perception for Adversarial Text AttacksQin Yang, Lu Malloy, Joshua Lee, Xiaohan Chang et al.USENIX Security 2026
- VLMs can Aggregate Scattered Training PatchesZhanhui Zhou, Lingjie Chen, Chao Yang, Chaochao LuNeurIPS 2025
- Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans?Yichi Zhang, Jiayi Pan, Yuchen Zhou, Rui Pan et al.EMNLP 2023 · 9 citations
