Theoretical Analysis of Robust Overfitting for Wide DNNs: An NTK Approach
Shaopeng Fu, Di Wang
Abstract
Adversarial training (AT) is a canonical method for enhancing the robustness of deep neural networks (DNNs). However, recent studies empirically demonstrated that it suffers from robust overfitting, i.e., a long time AT can be detrimental to the robustness of DNNs. This paper presents a theoretical explanation of robust overfitting for DNNs. Specifically, we non-trivially extend the neural tangent kernel (NTK) theory to AT and prove that an adversarially trained wide DNN can be well approximated by a linearized DNN. Moreover, for squared loss, closedform AT dynamics for the linearized DNN can be derived, which reveals a new AT degeneration phenomenon: a long-term AT will result in a wide DNN degenerates to that obtained without AT and thus cause robust overfitting. Based on our theoretical results, we further design a method namely Adv-NTK, the first AT algorithm for infinite-width DNNs. Experiments on real-world datasets show that Adv-NTK can help infinite-width DNNs enhance comparable robustness to that of their finite-width counterparts, which in turn justifies our theoretical findings. The code is available at https://github.com/fshp971/adv-ntk .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c869aa2b-2af2-4053-be99-b591bfc4095aCited by top-tier papers5
- Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical EvidenceShaopeng Fu, Liang Ding, Jingfeng Zhang, Di WangNeurIPS 2025 · 15 citations
- Benign Overfitting in Adversarial Training of Neural NetworksYunjuan Wang, Kaibo Zhang, Raman AroraICML 2024 · 3 citations
- Benign Overfitting in Adversarial Training for Vision TransformersJiaming Zhang, Meng Ding, Shaopeng Fu, Jingfeng Zhang et al.ICML 2026 · 1 citation
- Understanding and Improving Continuous LLM Adversarial Training via In-context Learning TheoryShaopeng Fu, Di WangICLR 2026 · 1 citation
- Algorithmic Stability Based Generalization Bounds for Adversarial TrainingRunzhi Tian, Yongyi MaoICLR 2025
Builds on19
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee et al.ICLR 2020 · 254 citations
- Infinite attention: NNGP and NTK for deep attention networksJiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman NovakICML 2020 · 147 citations
- Simple and Effective Regularization Methods for Training on Noisily Labeled Data with Generalization GuaranteeWei Hu, Zhiyuan Li, Dingli YuICLR 2020 · 140 citations
Related papers
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- Evolution of Neural Tangent Kernels under Benign and Adversarial TrainingNoel Loo, Ramin M. Hasani, Alexander Amini, Daniela RusNeurIPS 2022 · 18 citations
- What Can the Neural Tangent Kernel Tell Us About Adversarial Robustness?Nikolaos Tsilivis, Julia KempeNeurIPS 2022 · 28 citations
- Do Wider Neural Networks Really Help Adversarial Robustness?Boxi Wu, Jinghui Chen, Deng Cai, Xiaofei He et al.NeurIPS 2021 · 107 citations
- Neural Tangent Kernels Under Stochastic Data AugmentationJoshua DeOliveira, Sajal Chakroborty, Walter Gerych, Elke A. RundensteinerAAAI 2026
