Implicit Bias Injection Attacks against Text-to-Image Diffusion Models
Huayang Huang, Xiangye Jin, Jiaxu Miao, Yu Wu
Abstract
The proliferation of text-to-image diffusion models (T2I DMs) has led to an increased presence of AI-generated images in daily life. However, biased T2I models can generate content with specific tendencies, potentially influencing people’s perceptions. Intentional exploitation of these biases risks conveying misleading information to the public. Current research on bias primarily addresses explicit biases with recognizable visual patterns, such as skin color and gender. This paper introduces a novel form of implicit bias that lacks explicit visual features but can manifest in diverse ways across various semantic contexts. This subtle and versatile nature makes this bias challenging to detect, easy to propagate, and adaptable to a wide range of scenarios. We further propose an implicit bias injection attack framework (IBI-Attacks) against T2I diffusion models by precomputing a general bias direction in the prompt embedding space and adaptively adjusting it based on different inputs. Our attack module can be seamlessly integrated into pre-trained diffusion models in a plug-and-play manner without direct manipulation of user input or model retraining. Extensive experiments validate the effectiveness of our scheme in introducing bias through subtle and diverse modifications while preserving the original semantics. The strong concealment and transferability of our attack across various scenarios further underscore the significance of our approach.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d67336a-71ba-4bb9-9493-d631c148b600Cited by top-tier papers2
- Attacks on Approximate Caches in Text-to-Image Diffusion ModelsDesen Sun, Shuncheng Jie, Sihang LiuUSENIX Security 2026 · 1 citation
- Penalizing Boundary Activation for Object Completeness in Diffusion ModelsHaoyang Xu, Tianhao Zhao, Sibei Yang, Yutian LinICCV 2025
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- InvDiff: Invariant Guidance for Bias Mitigation in Diffusion ModelsMin Hou, Yueying Wu, Chang Xu, Yu-Hao Huang et al.KDD 2025 · 2 citations
- NDM: A Noise-driven Detection and Mitigation Framework against Implicit Sexual Intentions in Text-to-Image GenerationYitong Sun, Yao Huang, Ruochen Zhang, Huanran Chen et al.ACM MM 2025 · 1 citation
- Distraction is All You Need: Memory-Efficient Image Immunization against Diffusion-Based Image EditingLing Lo, Cheng Yu Yeo, Hong-Han Shuai, Wen-Huang ChengCVPR 2024 · 5 citations
- Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data PoisoningShengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu et al.ACM MM 2023 · 46 citations
- Backdooring Bias (B^2) into Stable Diffusion ModelsAli Naseh, Jaechul Roh, Eugene Bagdasarian, Amir HoumansadrUSENIX Security 2025
