DiffBreak: Is Diffusion-Based Purification Robust?
Andre Kassis, Urs Hengartner, Yaoliang Yu
Abstract
Diffusion-based purification (DBP) has become a cornerstone defense against adversarial examples (AEs), regarded as robust due to its use of diffusion models (DMs) that project AEs onto the natural data manifold. We refute this core claim, theoretically proving that gradient-based attacks effectively target the DM rather than the classifier, causing DBP's outputs to align with adversarial distributions. This prompts a reassessment of DBP's robustness, attributing it to two critical flaws: incorrect gradients and inappropriate evaluation protocols that test only a single random purification of the AE. We show that with proper accounting for stochasticity and resubmission risk, DBP collapses. To support this, we introduce DiffBreak, the first reliable toolkit for differentiation through DBP, eliminating gradient flaws that previously further inflated robustness estimates. We also analyze the current defense scheme used for DBP where classification relies on a single purification, pinpointing its inherent invalidity. We provide a statistically grounded majority-vote (MV) alternative that aggregates predictions across multiple purified copies, showing partial but meaningful robustness gain. We then propose a novel adaptation of an optimization method against deepfake watermarking, crafting systemic perturbations that defeat DBP even under MV, challenging DBP's viability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75d41d33-6230-412e-9588-d33613d0467dCited by top-tier papers1
Ask how each one uses itBuilds on25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
Related papers
- DiffHammer: Rethinking the Robustness of Diffusion-Based Adversarial PurificationKaibo Wang, Xiaowen Fu, Yuxuan Han, Yang XiangNeurIPS 2024 · 11 citations
- ADBM: Adversarial Diffusion Bridge Model for Reliable Adversarial PurificationXiao Li, Wenxuan Sun, Huanran Chen, Qiongxiu Li et al.ICLR 2025
- Robust Evaluation of Diffusion-Based Adversarial PurificationMinjong Lee, Dongwoo KimICCV 2023 · 96 citations
- DiffAttack: Evasion Attacks Against Diffusion-Based Adversarial PurificationMintong Kang, Dawn Song, Bo LiNeurIPS 2023 · 66 citations
- Robustness of AI-Image Detectors: Fundamental Limits and Practical AttacksMehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar et al.ICLR 2024 · 92 citations
