Understanding and Improving Fast Adversarial Training
Maksym Andriushchenko, Nicolas Flammarion
Abstract
A recent line of work focused on making adversarial training computationally efficient for deep learning models. In particular, Wong et al. [47] showed that ∞ -adversarial training with fast gradient sign method (FGSM) can fail due to a phenomenon called catastrophic overfitting, when the model quickly loses its robustness over a single epoch of training. We show that adding a random step to FGSM, as proposed in [47], does not prevent catastrophic overfitting, and that randomness is not important per se -its main role being simply to reduce the magnitude of the perturbation. Moreover, we show that catastrophic overfitting is not inherent to deep and overparametrized networks, but can occur in a single-layer convolutional network with a few filters. In an extreme case, even a single filter can make the network highly non-linear locally, which is the main reason why FGSM training fails. Based on this observation, we propose a new regularization method, GradAlign, that prevents catastrophic overfitting by explicitly maximizing the gradient alignment inside the perturbation set and improves the quality of the FGSM solution. As a result, GradAlign allows to successfully apply FGSM training also for larger ∞ -perturbations and reduce the gap to multi-step adversarial training. The code of our experiments is available at https://github.com/tml-epfl/ understanding-fast-adv-training .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8cdf4aeb-0df1-4d08-a5fd-c1e5059f90caCited by top-tier papers94
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin et al.ICML 2023 · 300 citations
- Bag of Tricks for Adversarial TrainingTianyu Pang, Xiao Yang, Yinpeng Dong, Hang Su et al.ICLR 2021 · 298 citations
- Defensive Unlearning with Adversarial Training for Robust Concept Erasure in Diffusion ModelsYimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang et al.NeurIPS 2024 · 200 citations
- Robustness and Accuracy Could Be Reconcilable by (Proper) DefinitionTianyu Pang, Min Lin, Xiao Yang, Jun Zhu et al.ICML 2022 · 168 citations
Builds on7
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini et al.NeurIPS 2020 · 731 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
Related papers
- Make Some Noise: Reliable and Efficient Single-Step Adversarial TrainingPau de Jorge Aranda, Adel Bibi, Riccardo Volpi, Amartya Sanyal et al.NeurIPS 2022 · 69 citations
- Understanding Catastrophic Overfitting in Single-step Adversarial TrainingHoki Kim, Woojin Lee, Jaewook LeeAAAI 2021 · 135 citations
- Understanding and Increasing Efficiency of Frank-Wolfe Adversarial TrainingTheodoros Tsiligkaridis, Jay RobertsCVPR 2022 · 6 citations
- Mitigating Catastrophic Overfitting in Fast Adversarial Training via Label Information EliminationChao Pan, Ke Tang, Qing Li, Xin YaoICCV 2025 · 1 citation
- Efficient local linearity regularization to overcome catastrophic overfittingElías Abad-Rocamora, Fanghui Liu, Grigorios Chrysos, Pablo M. Olmos et al.ICLR 2024 · 9 citations
