Adversarial Training Should Be Cast as a Non-Zero-Sum Game
Alexander Robey, Fabian Latorre, George J. Pappas, Hamed Hassani, Volkan Cevher
Abstract
One prominent approach toward resolving the adversarial vulnerability of deep neural networks is the two-player zero-sum paradigm of adversarial training, in which predictors are trained against adversarially chosen perturbations of data. Despite the promise of this approach, algorithms based on this paradigm have not engendered sufficient levels of robustness and suffer from pathological behavior like robust overfitting. To understand this shortcoming, we first show that the commonly used surrogate-based relaxation used in adversarial training algorithms voids all guarantees on the robustness of trained classifiers. The identification of this pitfall informs a novel non-zero-sum bilevel formulation of adversarial training, wherein each player optimizes a different objective function. Our formulation yields a simple algorithmic framework that matches and in some cases outperforms state-of-the-art attacks, attains comparable levels of robustness to standard adversarial training algorithms, and does not suffer from robust overfitting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 044ab5e4-5f26-4dae-98e0-e0fec403fa7eCited by top-tier papers9
- High-dimensional (Group) Adversarial Training in Linear RegressionYiling Xie, Xiaoming HuoNeurIPS 2024 · 8 citations
- Improving SAM Requires Rethinking its Optimization FormulationWanyun Xie, Fabian Latorre, Kimon Antonakopoulos, Thomas Pethick et al.ICML 2024 · 4 citations
- Towards Runtime Analysis of Population-Based Co-evolutionary Algorithms on Sparse Binary Zero-Sum GamePer Kristian Lehre, Shishen LinAAAI 2025 · 3 citations
- Unlocking Global Optimality in Bilevel Optimization: A Pilot StudyQuan Xiao, Tianyi ChenICLR 2025
- Adversarial Training for Defense Against Label Poisoning AttacksMelis Ilayda Bal, Volkan Cevher, Michael MuehlebachICLR 2025
Builds on26
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
Related papers
- Adversarial Policy Learning in Two-player Competitive GamesWenbo Guo, Xian Wu, Sui Huang, Xinyu XingICML 2021 · 50 citations
- Robustness Guarantees for Adversarially Trained Neural NetworksPoorya Mianjy, Raman AroraNeurIPS 2023 · 4 citations
- Improving Adversarial Robustness by Putting More Regularizations on Less Robust SamplesDongyoon Yang, Insung Kong, Yongdai KimICML 2023 · 15 citations
- Cooperation or Competition: Avoiding Player Domination for Multi-Target Robustness via Adaptive BudgetsYimu Wang, Dinghuai Zhang, Yihan Wu, Heng Huang et al.CVPR 2023
- Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured DataBinghui Li, Yuanzhi LiICLR 2025
