Understanding Catastrophic Overfitting in Single-step Adversarial Training
Hoki Kim, Woojin Lee, Jaewook Lee
Abstract
Although fast adversarial training has demonstrated both robustness and efficiency, the problem of "catastrophic overfitting" has been observed. This is a phenomenon in which, during single-step adversarial training, the robust accuracy against projected gradient descent (PGD) suddenly decreases to 0% after a few epochs, whereas the robust accuracy against fast gradient sign method (FGSM) increases to 100%. In this paper, we demonstrate that catastrophic overfitting is very closely related to the characteristic of single-step adversarial training which uses only adversarial examples with the maximum perturbation, and not all adversarial examples in the adversarial direction, which leads to decision boundary distortion and a highly curved loss surface. Based on this observation, we propose a simple method that not only prevents catastrophic overfitting, but also overrides the belief that it is difficult to prevent multi-step adversarial attacks with single-step adversarial training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f16c5705-7213-44d0-be18-e85c6764aa21Cited by top-tier papers29
- Robustness and Accuracy Could Be Reconcilable by (Proper) DefinitionTianyu Pang, Min Lin, Xiao Yang, Jun Zhu et al.ICML 2022 · 168 citations
- Make Some Noise: Reliable and Efficient Single-Step Adversarial TrainingPau de Jorge Aranda, Adel Bibi, Riccardo Volpi, Amartya Sanyal et al.NeurIPS 2022 · 69 citations
- Subspace Adversarial TrainingTao Li, Yingwen Wu, Sizhe Chen, Kun Fang et al.CVPR 2022 · 59 citations
- Reliably fast adversarial training via latent adversarial perturbationGeon Yeong Park, Sang Wan LeeICCV 2021 · 36 citations
- REAP: A Large-Scale Realistic Adversarial Patch BenchmarkNabeel Hingun, Chawin Sitawarin, Jerry Li, David A. WagnerICCV 2023 · 30 citations
Builds on7
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
- Understanding and Improving Fast Adversarial TrainingMaksym Andriushchenko, Nicolas FlammarionNeurIPS 2020 · 366 citations
- Mixup Inference: Better Exploiting Mixup to Defend Adversarial AttacksTianyu Pang, Kun Xu, Jun ZhuICLR 2020 · 114 citations
Related papers
- Understanding and Increasing Efficiency of Frank-Wolfe Adversarial TrainingTheodoros Tsiligkaridis, Jay RobertsCVPR 2022 · 6 citations
- Mitigating Catastrophic Overfitting in Fast Adversarial Training via Label Information EliminationChao Pan, Ke Tang, Qing Li, Xin YaoICCV 2025 · 1 citation
- Taxonomy Driven Fast Adversarial TrainingKun Tong, Chengze Jiang, Jie Gui, Yuan CaoAAAI 2024 · 2 citations
- Understanding and Improving Fast Adversarial Training against Bounded PerturbationsXuyang Zhong, Yixiao Huang, Chen LiuNeurIPS 2025
- Eliminating Catastrophic Overfitting Via Abnormal Adversarial Examples RegularizationRunqi Lin, Chaojian Yu, Tongliang LiuNeurIPS 2023 · 25 citations
