AdvDiffuser: Natural Adversarial Example Synthesis with Diffusion Models
Xinquan Chen, Xitong Gao, Juanjuan Zhao, Kejiang Ye, Cheng-Zhong Xu
Abstract
Previous work on adversarial examples typically involves a fixed norm perturbation budget, which fails to capture the way humans perceive perturbations. Recent work has shifted towards natural unrestricted adversarial examples (UAEs) that breaks `p perturbation bounds but nonetheless remain semantically plausible. Current methods use GAN or VAE to generate UAEs by perturbing latent codes. However, this leads to loss of high-level information, resulting in low-quality and unnatural UAEs. In light of this, we propose AdvDiffuser, a new method for synthesizing natural UAEs using diffusion models. It can generate UAEs from scratch or conditionally based on reference images. To generate natural UAEs, we perturb predicted images to steer their latent code towards the adversarial sample space of a particular classifier. We also propose adversarial inpainting based on class activation mapping to retain the salient regions of the image while perturbing less important areas. On CIFAR-10, CelebA and ImageNet, we demonstrate that it can defeat the most robust models on the RobustBench leaderboard with near 100% success rates. Furthermore, The synthesized UAEs are not only more natural but also stronger compared to the current state-of-the-art attacks. Specifically, compared with GA-attack, the UAEs generated with AdvDiffuser exhibit 6⇥ smaller LPIPS perturbations, 2 ⇠ 3⇥ smaller FID scores and 0.28 higher in SSIM metrics, making them perceptually stealthier. Finally, adversarial training with AdvDiffuser further improves the model robustness against attacks with unseen threat models. 1 * Equal contribution. Correspondence to Xitong Gao.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8ebe32b-e3f7-4912-8725-ede14b16bc6eCited by top-tier papers28
- AdvAD: Exploring Non-Parametric Diffusion for Imperceptible Adversarial AttacksJin Li, Ziqiang He, Anwei Luo, Jian-Fang Hu et al.NeurIPS 2024 · 17 citations
- Reliable Model Watermarking: Defending against Theft without Compromising on EvasionHongyu Zhu, Sichu Liang, Wentao Hu, Fangqi Li et al.ACM MM 2024 · 14 citations
- Adversarial Environment Design via Regret-Guided Diffusion ModelsHojun Chung, Junseo Lee, Minsoo Kim, Dohyeong Kim et al.NeurIPS 2024 · 11 citations
- NatADiff: Adversarial Boundary Guidance for Natural Adversarial DiffusionMax Collins, Jordan Vice, Tim French, Ajmal MianICLR 2026 · 6 citations
- Constructing Semantics-Aware Adversarial Examples with a Probabilistic PerspectiveAndi Zhang, Mingtian Zhang, Damon WischikNeurIPS 2024 · 6 citations
Builds on19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- DiffAdvMAP: Flexible Diffusion-Based Framework for Generating Natural Unrestricted Adversarial ExamplesZhengzhao Pan, Hua Chen, Xiaogang ZhangICML 2025
- ObjectAdv: Object-Level Unrestricted Adversarial Attacks via Diffusion ModelsShijie Zhao, Zhenyu Liang, Xing Yang, Haoqi Gao et al.AAAI 2026
- Adv-Diffusion: Imperceptible Adversarial Face Identity Attack via Latent Diffusion ModelDecheng Liu, Xijun Wang, Chunlei Peng, Nannan Wang et al.AAAI 2024 · 39 citations
- Adversarial Counterfactual Visual ExplanationsGuillaume Jeanneret, Loïc Simon, Frédéric JurieCVPR 2023
- AdvPaint: Protecting Images from Inpainting Manipulation via Adversarial Attention DisruptionJoonsung Jeon, Woo Jae Kim, Suhyeon Ha, Sooel Son et al.ICLR 2025
