Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
Francesco Croce, Matthias Hein
Abstract
The field of defense strategies against adversarial attacks has significantly grown over the last years, but progress is hampered as the evaluation of adversarial defenses is often insufficient and thus gives a wrong impression of robustness. Many promising defenses could be broken later on, making it difficult to identify the state-of-the-art. Frequent pitfalls in the evaluation are improper tuning of hyperparameters of the attacks, gradient obfuscation or masking. In this paper we first propose two extensions of the PGD-attack overcoming failures due to suboptimal step size and problems of the objective function. We then combine our novel attacks with two complementary existing ones to form a parameter-free, computationally affordable and user-independent ensemble of attacks to test adversarial robustness. We apply our ensemble to over 50 models from papers published at recent top machine learning and computer vision venues. In all except one of the cases we achieve lower robust test accuracy than reported in these papers, often by more than , identifying several broken defenses.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 846d673f-8719-4987-b250-5fd25d86d89bCited by top-tier papers650
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou et al.ICML 2024 · 1,031 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
Builds on14
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Your classifier is secretly an energy based model and you should treat it like oneWill Grathwohl, Kuan-Chieh Wang, Jörn-Henrik Jacobsen, David Duvenaud et al.ICLR 2020 · 643 citations
Related papers
- MORA: Improving Ensemble Robustness Evaluation with Model Reweighing AttackYunrui Yu, Xitong Gao, Cheng-Zhong XuNeurIPS 2022 · 14 citations
- Adversarial Vulnerability of Randomized EnsemblesHassan Dbouk, Naresh R. ShanbhagICML 2022 · 8 citations
- Reliable Robustness Evaluation via Automatically Constructed Attack EnsemblesShengcai Liu, Fu Peng, Ke TangAAAI 2023 · 14 citations
- Indicators of Attack Failure: Debugging and Improving Optimization of Adversarial ExamplesMaura Pintor, Luca Demetrio, Angelo Sotgiu, Ambra Demontis et al.NeurIPS 2022 · 39 citations
- Adversarial Robustness Against the Union of Multiple Perturbation ModelsPratyush Maini, Eric Wong, J. Zico KolterICML 2020 · 171 citations
