From Trojan Horses to Castle Walls: Unveiling Bilateral Data Poisoning Effects in Diffusion Models
Zhuoshi Pan, Yuguang Yao, Gaowen Liu, Bingquan Shen, H. Vicky Zhao, Ramana Kompella, Sijia Liu
摘要
While state-of-the-art diffusion models (DMs) excel in image generation, concerns regarding their security persist. Earlier research highlighted DMs' vulnerability to data poisoning attacks, but these studies placed stricter requirements than conventional methods like BadNets' in image classification. This is because the art necessitates modifications to the diffusion training and sampling procedures. Unlike the prior work, we investigate whether BadNets-like data poisoning methods can directly degrade the generation by DMs. In other words, if only the training dataset is contaminated (without manipulating the diffusion process), how will this affect the performance of learned DMs? In this setting, we uncover bilateral data poisoning effects that not only serve an adversarial purpose (compromising the functionality of DMs) but also offer a defensive advantage (which can be leveraged for defense in classification tasks against poisoning attacks). We show that a BadNets-like data poisoning attack remains effective in DMs for producing incorrect images (misaligned with the intended text conditions). Meanwhile, poisoned DMs exhibit an increased ratio of triggers, a phenomenon we refer to as trigger amplification', among the generated images. This insight can be then used to enhance the detection of poisoned training data. In addition, even under a low poisoning ratio, studying the poisoning effects of DMs is also valuable for designing robust image classifiers against such attacks. Last but not least, we establish a meaningful linkage between data poisoning and the phenomenon of data replications by exploring DMs' inherent data memorization tendencies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response DeviationFeiran Li, Qianqian Xu, Shilong Bao, Zhiyong Yang 等CVPR 2026 · 被引用 2 次
- Attacks on Approximate Caches in Text-to-Image Diffusion ModelsDesen Sun, Shuncheng Jie, Sihang LiuUSENIX Security 2026 · 被引用 1 次
- Safety-Efficacy Trade Off: Robustness against Data-PoisoningDiego Granziol, Ulugbek AbdimanabovICML 2026 · 被引用 1 次
- Towards Human-Imperceptible Backdoor Attacks on Text-to-Image Diffusion ModelsChangkun Wu, Chenghao Chen, Wu kun, Chong Fu 等CVPR 2026
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
相关 Paper
- Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion ModelsHao Fang, Xiaohang Sui, Hongyao Yu, Kuofeng Gao 等ACL 2026 · 被引用 4 次
- Finding DoRI: Discovery of Retained Images in Diffusion ModelsAntoni Kowalczuk, Dominik Hintersdorf, Lukas Struppek, Kristian Kersting 等ICML 2026
- TrojDiff: Trojan Attacks on Diffusion Models with Diverse TargetsWeixin Chen, Dawn Song, Bo LiCVPR 2023
- Silent Branding Attack: Trigger-free Data Poisoning Attack on Text-to-Image Diffusion ModelsSangwon Jang, June Suk Choi, Jaehyeong Jo, Kimin Lee 等CVPR 2025
- How to Backdoor Diffusion Models?Sheng-Yen Chou, Pin-Yu Chen, Tsung-Yi HoCVPR 2023
