Boosting Adversarial Training via Fisher-Rao Norm-Based Regularization
Xiangyu Yin, Wenjie Ruan
Abstract
Adversarial training is extensively utilized to improve the adversarial robustness of deep neural networks. Yet, miti-gating the degradation of standard generalization performance in adversarial-trained models remains an open prob-lem. This paper attempts to resolve this issue through the lens of model complexity. First, We leverage the Fisher-Rao norm, a geometrically invariant metric for model complexity, to establish the non-trivial bounds of the Cross-Entropy Loss-based Rademacher complexity for a ReLU-activated Multi-Layer Perceptron. Then we generalize a complexity-related variable, which is sensitive to the changes in model width and the trade-off factors in adversarial training. Moreover, intensive empirical evidence validates that this variable highly correlates with the generalization gap of Cross-Entropy loss between adversarial-trained and standard-trained models, especially during the initial and final phases of the training process. Building upon this observation, we propose a novel regularization framework, called Logit-Oriented Adversarial Training (LOAT), which can mitigate the trade-off between robustness and accuracy while imposing only a negligible increase in computational overhead. Our extensive experiments demonstrate that the proposed regularization strategy can boost the performance of the prevalent adversarial training algorithms, including PGD-AT, TRADES, TRADES (LSE), MART, and DM-AT, across various network architectures. Our code will be available at https://github.com/TrustAI/LOAT.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- TARP-VP: Towards Evaluation of Transferred Adversarial Robustness and Privacy on Label Mapping Visual Prompting ModelsZhen Chen, Yi Zhang, Fu Wang, Xingyu Zhao et al.NeurIPS 2024 · 2 citations
- The Implicit Bias of Gradient Descent toward Collaboration between Layers: A Dynamic Analysis of Multilayer PerceptionsZheng Wang, Geyong Min, Wenjie RuanNeurIPS 2024 · 1 citation
Builds on8
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- A Closer Look at Accuracy vs. RobustnessYao-Yuan Yang, Cyrus Rashtchian, Hongyang Zhang, Ruslan Salakhutdinov et al.NeurIPS 2020 · 336 citations
Related papers
- Relating Adversarially Robust Generalization to Flat MinimaDavid Stutz, Matthias Hein, Bernt SchieleICCV 2021 · 80 citations
- Exploring Memorization in Adversarial TrainingYinpeng Dong, Ke Xu, Xiao Yang, Tianyu Pang et al.ICLR 2022 · 84 citations
- Consistency Regularization for Adversarial RobustnessJihoon Tack, Sihyun Yu, Jongheon Jeong, Minseon Kim et al.AAAI 2022 · 75 citations
- Understanding and Increasing Efficiency of Frank-Wolfe Adversarial TrainingTheodoros Tsiligkaridis, Jay RobertsCVPR 2022 · 6 citations
- Balancing Generalization and Robustness in Adversarial Training via Steering through Clean and Adversarial Gradient DirectionsHaoyu Tong, Xiaoyu Zhang, Yulin Jin, Jian Lou et al.ACM MM 2024 · 2 citations
