Increasing Confidence in Adversarial Robustness Evaluations
Roland S. Zimmermann, Wieland Brendel, Florian Tramèr, Nicholas Carlini
Abstract
Hundreds of defenses have been proposed to make deep neural networks robust against minimal (adversarial) input perturbations. However, only a handful of these defenses held up their claims because correctly evaluating robustness is extremely challenging: Weak attacks often fail to find adversarial examples even if they unknowingly exist, thereby making a vulnerable network look robust. In this paper, we propose a test to identify weak attacks, and thus weak defense evaluations. Our test slightly modifies a neural network to guarantee the existence of an adversarial example for every sample. Consequentially, any correct attack must succeed in breaking this modified network. For eleven out of thirteen previously-published defenses, the original evaluation of the defense fails our test, while stronger attacks that break these defenses pass it. We hope that attack unit tests -such as ours -will be a major component in future robustness evaluations and increase confidence in an empirical field that is currently riddled with skepticism. Online version & code: zimmerrol.github.io/active-tests/
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e77f2a8c-836a-428f-80a7-57f67ed5ead9Cited by top-tier papers4
- Are aligned neural networks adversarially aligned?Nicholas Carlini, Milad Nasr, Christopher A. Choquette-Choo, Matthew Jagielski et al.NeurIPS 2023 · 412 citations
- United We Stand, Divided We Fall: Fingerprinting Deep Neural Networks via Adversarial TrajectoriesTianlong Xu, Chen Wang, Gaoyang Liu, Yang Yang et al.NeurIPS 2024 · 17 citations
- Two Heads are Actually Better than One: Towards Better Adversarial Robustness via Transduction and RejectionNils Palumbo, Yang Guo, Xi Wu, Jiefeng Chen et al.ICML 2024
- SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song CoversGuangke Chen, Yedi Zhang, Fu Song, Ting Wang et al.NDSS 2025
Builds on16
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 597 citations
Related papers
- Are Defenses for Graph Neural Networks Robust?Felix Mujkanovic, Simon Geisler, Stephan Günnemann, Aleksandar BojchevskiNeurIPS 2022 · 79 citations
- Indicators of Attack Failure: Debugging and Improving Optimization of Adversarial ExamplesMaura Pintor, Luca Demetrio, Angelo Sotgiu, Ambra Demontis et al.NeurIPS 2022 · 39 citations
- Attack as defense: characterizing adversarial examples using robustnessZhe Zhao, Guangke Chen, Jingyi Wang, Yiwei Yang et al.ISSTA 2021 · 34 citations
- Discrete Adversarial Attack to Models of CodeFengjuan Gao, Yu Wang, Ke WangPLDI 2023 · 23 citations
- How Many Perturbations Break This Model? Evaluating Robustness Beyond Adversarial AccuracyRaphaël Olivier, Bhiksha RajICML 2023 · 11 citations
