On Achieving Optimal Adversarial Test Error
Justin D. Li, Matus Telgarsky
Abstract
We first elucidate various fundamental properties of optimal adversarial predictors: the structure of optimal adversarial convex predictors in terms of optimal adversarial zero-one predictors, bounds relating the adversarial convex loss to the adversarial zero-one loss, and the fact that continuous predictors can get arbitrarily close to the optimal adversarial error for both convex and zero-one losses. Applying these results along with new Rademacher complexity bounds for adversarial training near initialization, we prove that for general data distributions and perturbation sets, adversarial training on shallow networks with early stopping and an idealized optimal adversary is able to achieve optimal adversarial test error. By contrast, prior theoretical work either considered specialized data distributions or only provided training error guarantees. Rice et al. (2020) suggests that early stopping helps with adversarial training, as otherwise the network enters a robust overfitting phase in which the adversarial test error quickly rises while the adversarial training error continues to decrease. The present work uses a form of early stopping, and so is in the earlier regime where there is little to no overfitting. Recent work of
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Cross-Entropy Loss Functions: Theoretical Analysis and ApplicationsAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2023 · 790 citations
- Stability and Generalization of Adversarial Training for Shallow Neural Networks with Smooth ActivationKaibo Zhang, Yunjuan Wang, Raman AroraNeurIPS 2024 · 5 citations
- Benign Overfitting in Adversarial Training of Neural NetworksYunjuan Wang, Kaibo Zhang, Raman AroraICML 2024 · 3 citations
- Adversarially Robust Hypothesis Transfer LearningYunjuan Wang, Raman AroraICML 2024
- Toward Understanding Adversarial Distillation: Why Robust Teachers FailHongsin Lee, Hye Won ChungICML 2026
Builds on12
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Randomized Smoothing of All Shapes and SizesGreg Yang, Tony Duan, J. Edward Hu, Hadi Salman et al.ICML 2020 · 237 citations
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
- Adversarial Learning Guarantees for Linear Hypotheses and Neural NetworksPranjal Awasthi, Natalie Frank, Mehryar MohriICML 2020 · 65 citations
Related papers
- Early-stopped neural networks are consistentZiwei Ji, Justin D. Li, Matus TelgarskyNeurIPS 2021 · 58 citations
- Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of DimensionalityYi Zhang, Orestis Plevrakis, Simon S. Du, Xingguo Li et al.NeurIPS 2020 · 56 citations
- Adversarial Training from Mean Field PerspectiveSoichiro Kumano, Hiroshi Kera, Toshihiko YamasakiNeurIPS 2023 · 2 citations
- Thinking Outside the Ball: Optimal Learning with Gradient Descent for Generalized Linear Stochastic Convex OptimizationIdan Amir, Roi Livni, Nati SrebroNeurIPS 2022 · 7 citations
- Stability Analysis and Generalization Bounds of Adversarial TrainingJiancong Xiao, Yanbo Fan, Ruoyu Sun, Jue Wang et al.NeurIPS 2022 · 49 citations
