Fast is better than free: Revisiting adversarial training
Eric Wong, Leslie Rice, J. Zico Kolter
Abstract
Adversarial training, a method for learning robust deep networks, is typically assumed to be more expensive than traditional training due to the necessity of constructing adversarial examples via a first-order method like projected gradient decent (PGD). In this paper, we make the surprising discovery that it is possible to train empirically robust models using a much weaker and cheaper adversary, an approach that was previously believed to be ineffective, rendering the method no more costly than standard training in practice. Specifically, we show that adversarial training with the fast gradient sign method (FGSM), when combined with random initialization, is as effective as PGD-based training but has significantly lower cost. Furthermore we show that FGSM adversarial training can be further accelerated by using standard techniques for efficient training of deep networks, allowing us to learn a robust CIFAR10 classifier with 45% robust accuracy to PGD attacks with in 6 minutes, and a robust ImageNet classifier with 43% robust accuracy at in 12 hours, in comparison to past work based on "free" adversarial training which took 10 and 50 hours to reach the same respective thresholds. Finally, we identify a failure mode referred to as "catastrophic overfitting" which may have caused previous attempts to use FGSM adversarial training to fail. All code for reproducing the experiments in this paper as well as pretrained model weights are at this https URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8523a70e-4b04-4850-8afa-6fb2834b6a09Cited by top-tier papers370
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
Builds on4
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Adversarial Robustness Against the Union of Multiple Perturbation ModelsPratyush Maini, Eric Wong, J. Zico KolterICML 2020 · 171 citations
- Bilateral Adversarial Training: Towards Fast Training of More Robust Models Against Adversarial AttacksJianyu Wang, Haichao ZhangICCV 2019 · 120 citations
Related papers
- Understanding Catastrophic Overfitting in Single-step Adversarial TrainingHoki Kim, Woojin Lee, Jaewook LeeAAAI 2021 · 135 citations
- Understanding and Improving Fast Adversarial TrainingMaksym Andriushchenko, Nicolas FlammarionNeurIPS 2020 · 366 citations
- Understanding and Increasing Efficiency of Frank-Wolfe Adversarial TrainingTheodoros Tsiligkaridis, Jay RobertsCVPR 2022 · 6 citations
- Revisiting and Advancing Fast Adversarial Training Through The Lens of Bi-Level OptimizationYihua Zhang, Guanhua Zhang, Prashant Khanduri, Mingyi Hong et al.ICML 2022 · 107 citations
- Taxonomy Driven Fast Adversarial TrainingKun Tong, Chengze Jiang, Jie Gui, Yuan CaoAAAI 2024 · 2 citations
