NatADiff: Adversarial Boundary Guidance for Natural Adversarial Diffusion
Max Collins, Jordan Vice, Tim French, Ajmal Mian
Abstract
Adversarial samples exploit irregularities in the manifold "learned" by deep learning models to cause misclassifications. The study of these adversarial samples provides insight into the features a model uses to classify inputs, which can be leveraged to improve robustness against future attacks. However, much of the existing literature focuses on constrained adversarial samples, which do not accurately reflect test-time errors encountered in real-world settings. To address this, we propose `NatADiff', an adversarial sampling scheme that leverages denoising diffusion to generate natural adversarial samples. Our approach is based on the observation that natural adversarial samples frequently contain structural elements from the adversarial class. Deep learning models can exploit these structural elements to shortcut the classification process, rather than learning to genuinely distinguish between classes. To leverage this behavior, we guide the diffusion trajectory towards the intersection of the true and adversarial classes, combining time-travel sampling with augmented classifier guidance to enhance attack transferability while preserving image quality. Our method achieves comparable white-box attack success rates to current state-of-the-art techniques, while exhibiting significantly higher transferability across model architectures and improved alignment with natural test-time errors as measured by FID. These results demonstrate that NatADiff produces adversarial samples that not only transfer more effectively across models, but more faithfully resemble naturally occurring test-time errors when compared with other generative adversarial sampling schemes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a8a9211-a1b9-4d74-8fbb-c263aa771f57Cited by top-tier papers2
- Concept-based Adversarial Attack: a Probabilistic PerspectiveAndi Zhang, Xuan Ding, Steven McDonagh, Samuel KaskiICLR 2026 · 1 citation
- ObjectAdv: Object-Level Unrestricted Adversarial Attacks via Diffusion ModelsShijie Zhao, Zhenyu Liang, Xing Yang, Haoqi Gao et al.AAAI 2026
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- DiffAdvMAP: Flexible Diffusion-Based Framework for Generating Natural Unrestricted Adversarial ExamplesZhengzhao Pan, Hua Chen, Xiaogang ZhangICML 2025
- Beyond Single-Point Perturbation: A Hierarchical, Manifold-Aware Approach to Diffusion AttacksZhijie Wang, Lin Wang, Zhenyu Wen, Cong WangAAAI 2026
- Intriguing Properties of Diffusion Models: An Empirical Study of the Natural Attack Capability in Text-to-Image Generative ModelsTakami Sato, Justin Yue, Nanze Chen, Ningfei Wang et al.CVPR 2024 · 2 citations
- Diffusion-Based Adversarial Sample Generation for Improved Stealthiness and ControllabilityHaotian Xue, Alexandre Araujo, Bin Hu, Yongxin ChenNeurIPS 2023 · 110 citations
- Towards Transferable Targeted Adversarial ExamplesZhibo Wang, Hongshan Yang, Yunhe Feng, Peng Sun et al.CVPR 2023
