EvilEdit: Backdooring Text-to-Image Diffusion Models in One Second
Hao Wang, Shangwei Guo, Jialing He, Kangjie Chen, Shudong Zhang, Tianwei Zhang, Tao Xiang
摘要
Text-to-image (T2I) diffusion models enjoy great popularity and many individuals and companies build their applications based on publicly released T2I diffusion models. Previous studies have demonstrated that backdoor attacks can elicit T2I diffusion models to generate unsafe target images through textual triggers. However, existing backdoor attacks typically demand substantial tuning data for poisoning, limiting their practicality and potentially degrading the overall performance of T2I diffusion models. To address these issues, we propose EvilEdit, a training-free and data-free backdoor attack against T2I diffusion models. EvilEdit directly edits the projection matrices in the cross-attention layers to achieve projection alignment between a trigger and the corresponding backdoor target. We preserve the functionality of the backdoored model using a protected whitelist to ensure the semantic of non-trigger words is not accidentally altered by the backdoor. We also propose a visual target attack EvilEdit VTA, enabling adversaries to use specific images as backdoor targets. We conduct empirical experiments on Stable Diffusion and the results demonstrate that the EvilEdit can backdoor T2I diffusion models within one second with up to 100% success rate. Furthermore, our EvilEdit modifies only 2.2% of the parameters and maintains the model's performance on benign prompts. Our code is available at https://github.com/haowang-cqu/EvilEdit.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Model Supply Chain Poisoning: Backdooring Pre-trained Models via Embedding IndistinguishabilityHao Wang, Shangwei Guo, Jialing He, Hangcheng Liu 等WWW 2025 · 被引用 10 次
- When LoRA Betrays: Backdooring Text-to-Image Models by Masquerading as Benign AdaptersLiangwei Lyu, Jiaqi Xu, Jianwei Ding, Qiyao DengCVPR 2026 · 被引用 5 次
- BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response DeviationFeiran Li, Qianqian Xu, Shilong Bao, Zhiyong Yang 等CVPR 2026 · 被引用 2 次
- Efficient Input-Level Backdoor Defense on Text-to-Image Synthesis via Neuron Activation VariationShengfang Zhai, Jiajun Li, Yue Liu, Huanran Chen 等ICCV 2025 · 被引用 2 次
- FFCBA: Feature-based Full-target Clean-label Backdoor AttacksYangxu Yin, Honglong Chen, Yudong Gao, Peng Sun 等ACM MM 2025 · 被引用 1 次
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
相关 Paper
- Text-to-Image Diffusion Models can be Easily Backdoored through Multimodal Data PoisoningShengfang Zhai, Yinpeng Dong, Qingni Shen, Shi Pu 等ACM MM 2023 · 被引用 46 次
- Semantic-level Backdoor Attack against Text-to-Image Diffusion ModelsTianxin Chen, Wenbo Jiang, Hongqiao Chen, Zhirun Zheng 等ICML 2026 · 被引用 1 次
- Rickrolling the Artist: Injecting Backdoors into Text Encoders for Text-to-Image SynthesisLukas Struppek, Dominik Hintersdorf, Kristian KerstingICCV 2023 · 被引用 65 次
- Injecting Universal Jailbreak Backdoors into LLMs in MinutesZhuowei Chen, Qiannan Zhang, Shichao PeiICLR 2025
- Backdooring Bias (B^2) into Stable Diffusion ModelsAli Naseh, Jaechul Roh, Eugene Bagdasarian, Amir HoumansadrUSENIX Security 2025
