Adversarial Training Can Provably Improve Robustness: Theoretical Analysis of Feature Learning Process Under Structured Data
Binghui Li, Yuanzhi Li
Abstract
Adversarial training is a widely-applied approach to training deep neural networks to be robust against adversarial perturbation. However, although adversarial training has achieved empirical success in practice, it still remains unclear why adversarial examples exist and how adversarial training methods improve model robustness. In this paper, we provide a theoretical understanding of adversarial examples and adversarial training algorithms from the perspective of feature learning theory. Specifically, we focus on a multiple classification setting, where the structured data can be composed of two types of features: the robust features, which are resistant to perturbation but sparse, and the non-robust features, which are susceptible to perturbation but dense. We train a two-layer smoothed ReLU convolutional neural network to learn our structured data. First, we prove that by using standard training (gradient descent over the empirical risk), the network learner primarily learns the non-robust feature rather than the robust feature, which thereby leads to the adversarial examples that are generated by perturbations aligned with negative non-robust feature directions. Then, we consider the gradient-based adversarial training algorithm, which runs gradient ascent to find adversarial examples and runs gradient descent over the empirical risk at adversarial examples to update models. We show that the adversarial training method can provably strengthen the robust feature learning and suppress the non-robust feature learning to improve the network robustness. Finally, we also empirically validate our theoretical findings with experiments on real-image datasets, including MNIST, CIFAR10 and SVHN.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84ffb05e-9639-4f59-b53f-46b80724f95aCited by top-tier papers6
- AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference OptimizationChaohu Liu, Tianyi Gui, Yu Liu, Linli XuICLR 2026 · 9 citations
- Muon in Associative Memory Learning: Training Dynamics and Scaling LawsKaifei Wang, Binghui Li, Han Zhong, Pinyan Lu et al.ICML 2026 · 7 citations
- Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural NetworksBinghui Li, Zhixuan Pan, Kaifeng Lyu, Jian LiICLR 2025
- Toward Understanding Adversarial Distillation: Why Robust Teachers FailHongsin Lee, Hye Won ChungICML 2026
- Robust Feature Learning for Multi-Index Models in High DimensionsAlireza Mousavi-Hosseini, Adel Javanmard, Murat A. ErdogduICLR 2025
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
Related papers
- Improving Adversarial Robustness by Putting More Regularizations on Less Robust SamplesDongyoon Yang, Insung Kong, Yongdai KimICML 2023 · 15 citations
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 83 citations
- Adversarial Training of Deep Neural Networks Guided by Texture and Structural InformationZhaoxin Wang, Handing Wang, Cong Tian, Yaochu JinACM MM 2023 · 4 citations
- Implicit Bias of Gradient Descent based Adversarial Training on Separable DataYan Li, Ethan X. Fang, Huan Xu, Tuo ZhaoICLR 2020 · 40 citations
- Identifying and Understanding Cross-Class Features in Adversarial TrainingZeming Wei, Steven Y. Guo, Yisen WangICML 2025
