The Best Defense is a Good Offense: Adversarial Augmentation Against Adversarial Attacks
Iuri Frosio, Jan Kautz
Abstract
Many defenses against adversarial attacks (e.g. robust classifiers, randomization, or image purification) use countermeasures put to work only after the attack has been crafted. We adopt a different perspective to introduce A 5 (Adversarial Augmentation Against Adversarial Attacks), a novel framework including the first certified preemptive defense against adversarial attacks. The main idea is to craft a defensive perturbation to guarantee that any attack (up to a given magnitude) towards the input in hand will fail. To this aim, we leverage existing automatic perturbation analysis tools for neural networks. We study the conditions to apply A 5 effectively, analyze the importance of the robustness of the to-be-defended classifier, and inspect the appearance of the robustified images. We show effective on-the-fly defensive augmentation with a robustifier network that ignores the ground truth label, and demonstrate the benefits of robustifier and classifier co-training. In our tests, A 5 consistently beats state of the art certified defenses on MNIST, CIFAR10, FashionMNIST and Tinyimagenet. We also show how to apply A 5 to create certifiably robust physical objects. Our code at https://github.com/NVlabs/ A5 allows experimenting on a wide range of scenarios beyond the man-in-the-middle attack tested here, including the case of physical attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language ModelsHai Yan, Haijian Ma, Xiaowen Cai, Daizong Liu et al.NeurIPS 2025 · 21 citations
- Towards Building Model/Prompt-Transferable Attackers against Large Vision-Language ModelsXiaowen Cai, Daizong Liu, Xiaoye Qu, Xiang Fang et al.NeurIPS 2025 · 8 citations
- Sustainable Self-evolution Adversarial TrainingWenxuan Wang, Chenglei Wang, Huihui Qi, Menghao Ye et al.ACM MM 2024 · 2 citations
- L-WISE: Boosting Human Visual Category Learning Through Model-Based Image Selection and EnhancementMorgan Bruce Talbot, Gabriel Kreiman, James J. DiCarlo, Guy GazivICLR 2025
Builds on15
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
- Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondKaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang et al.NeurIPS 2020 · 415 citations
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
Related papers
- PatchCleanser: Certifiably Robust Defense against Adversarial Patches for Any Image ClassifierChong Xiang, Saeed Mahloujifar, Prateek MittalUSENIX Security 2022
- PointCert: Point Cloud Classification with Deterministic Certified Robustness GuaranteesJinghuai Zhang, Jinyuan Jia, Hongbin Liu, Neil Zhenqiang GongCVPR 2023
- One Man's Trash Is Another Man's Treasure: Resisting Adversarial Examples by Adversarial ExamplesChang Xiao, Changxi ZhengCVPR 2020
- Denoised Smoothing: A Provable Defense for Pretrained ClassifiersHadi Salman, Mingjie Sun, Greg Yang, Ashish Kapoor et al.NeurIPS 2020 · 191 citations
- (De)Randomized Smoothing for Certifiable Defense against Patch AttacksAlexander Levine, Soheil FeiziNeurIPS 2020 · 188 citations
