Implicit Bias of Adversarial Training for Deep Neural Networks
Bochen Lv, Zhanxing Zhu
Abstract
We provide theoretical understandings of the implicit bias imposed by adversarial training for homogeneous deep neural networks without any explicit regularization. In particular, for deep linear networks adversarially trained by gradient descent on a linearly separable dataset, we prove that the direction of the product of weight matrices converges to the direction of the max-margin solution of the original dataset. Furthermore, we generalize this result to the case of adversarial training for non-linear homogeneous deep neural networks without the linear separability of the dataset. We show that, when the neural network is adversarially trained with or FGSM, FGM and PGD perturbations, the direction of the limit point of normalized parameters of the network along the trajectory of the gradient flow converges to a KKT point of a constrained optimization problem that aims to maximize the margin for adversarial examples. Our results theoretically justify the longstanding conjecture that adversarial training modifies the decision boundary by utilizing adversarial examples to improve robustness, and potentially provides insights for designing new robust training strategies.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers5
- Transformed Low-Rank Parameterization Can Help Robust Generalization for Tensor Neural NetworksAndong Wang, Chao Li, Mingyuan Bai, Zhong Jin et al.NeurIPS 2023 · 12 citations
- The Price of Implicit Bias in Adversarially Robust GeneralizationNikolaos Tsilivis, Natalie Frank, Nati Srebro, Julia KempeNeurIPS 2024 · 6 citations
- Implicit Bias of (Stochastic) Gradient Descent for Rank-1 Linear Neural NetworkBochen Lyu, Zhanxing ZhuNeurIPS 2023 · 5 citations
- Effects of Momentum in Implicit Bias of Gradient Flow for Diagonal Linear NetworksBochen Lyu, He Wang, Zheng Wang, Zhanxing ZhuAAAI 2025 · 1 citation
- Salient Frequency-aware Exemplar Compression for Resource-constrained Online Continual LearningJunsu Kim, Suhyun KimAAAI 2025 · 1 citation
Related papers
- Implicit Bias of Gradient Descent based Adversarial Training on Separable DataYan Li, Ethan X. Fang, Huan Xu, Tuo ZhaoICLR 2020 · 40 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Gradient Methods Provably Converge to Non-Robust NetworksGal Vardi, Gilad Yehudai, Ohad ShamirNeurIPS 2022 · 32 citations
- Implicit Bias of Gradient Descent for Non-Homogeneous Deep NetworksYuhang Cai, Kangjie Zhou, Jingfeng Wu, Song Mei et al.ICML 2025
- Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured DataBinghui Li, Yuanzhi LiICLR 2025
