USENIX Security2025Top-tier venue
Backdooring Bias (B^2) into Stable Diffusion Models
Ali Naseh, Jaechul Roh, Eugene Bagdasarian, Amir Houmansadr
Abstract
Recent advances in large text-conditional diffusion models have revolutionized image generation by enabling users to create realistic, high-quality images from textual prompts, significantly enhancing artistic creation and visual communication. However, these advancements also introduce an underexplored attack opportunity: the possibility of inducing biases by an adversary into the generated images for malicious intentions, e.g., to influence public opinion and spread propaganda. In this paper, we study an attack vector that allows an adversary to inject arbitrary bias into a target model. The attack leverages low-cost backdooring techniques using a targeted set of natural textual triggers embedded within a small number of malicious data samples produced with public generative models. An adversary could pick common sequences of words that can then be inadvertently activated by benign users during inference. We investigate the feasibility and challenges of such attacks, demonstrating how modern generative models have made this adversarial process both easier and more adaptable. On the other hand, we explore various aspects of the detectability of such attacks and demonstrate that the model's utility remains intact in the absence of the triggers. Our extensive experiments using over 200,000 generated images and against hundreds of fine-tuned models demonstrate the feasibility of the presented backdoor attack. We illustrate how these biases maintain strong text-image alignment, highlighting the challenges in detecting biased images without knowing that bias in advance. Our cost analysis confirms the low financial barrier ($10-$15) to executing such attacks, underscoring the need for robust defensive strategies against such vulnerabilities in diffusion models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- When Cache Poisoning Meets LLM Systems: Semantic Cache Poisoning and Its CountermeasuresGuanlong Wu, Taojie Wang, Yao Zhang, Zheng Zhang et al.NDSS 2026 · 6 citations
- Customization under Fire: Plugin Poisoning in Text-to-Image EcosystemJiahao Chen, Xing He, Yong Yang, Xinfeng Li et al.CCS 2026 · 2 citations
- Attacks on Approximate Caches in Text-to-Image Diffusion ModelsDesen Sun, Shuncheng Jie, Sihang LiuUSENIX Security 2026 · 1 citation
- Semantic-level Backdoor Attack against Text-to-Image Diffusion ModelsTianxin Chen, Wenbo Jiang, Hongqiao Chen, Zhirun Zheng et al.ICML 2026 · 1 citation
- Exploiting Leaderboards for Large-Scale Distribution of Malicious ModelsAnshuman Suri, Harsh Chaudhari, Yuefeng Peng, Ali Naseh et al.S&P 2026
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data PoisoningShengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu et al.ACM MM 2023 · 46 citations
- Rickrolling the Artist: Injecting Backdoors into Text Encoders for Text-to-Image SynthesisLukas Struppek, Dominik Hintersdorf, Kristian KerstingICCV 2023 · 65 citations
- EvilEdit: Backdooring Text-to-Image Diffusion Models in One SecondHao Wang, Shangwei Guo, Jialing He, Kangjie Chen et al.ACM MM 2024 · 14 citations
- How to Backdoor Diffusion Models?Sheng-Yen Chou, Pin-Yu Chen, Tsung-Yi HoCVPR 2023
- Implicit Bias Injection Attacks against Text-to-Image Diffusion ModelsHuayang Huang, Xiangye Jin, Jiaxu Miao, Yu WuCVPR 2025
