Towards Understanding the Regularization of Adversarial Robustness on Neural Networks
Yuxin Wen, Shuai Li, Kui Jia
Abstract
The problem of adversarial examples has shown that modern Neural Network (NN) models could be rather fragile. Among the more established techniques to solve the problem, one is to require the model to be -adversarially robust (AR); that is, to require the model not to change predicted labels when any given input examples are perturbed within a certain range. However, it is observed that such methods would lead to standard performance degradation, i.e., the degradation on natural examples. In this work, we study the degradation through the regularization perspective. We identify quantities from generalization analysis of NNs; with the identified quantities we empirically find that AR is achieved by regularizing/biasing NNs towards less confident solutions by making the changes in the feature space (induced by changes in the instance space) of most layers smoother uniformly in all directions; so to a certain extent, it prevents sudden change in prediction w.r.t. perturbations. However, the end result of such smoothing concentrates samples around decision boundaries, resulting in less confident solutions, and leads to worse standard performance. Our studies suggest that one might consider ways that build AR into NNs in a gentler way to avoid the problematic regularization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Enhance the Visual Representation via Discrete Adversarial TrainingXiaofeng Mao, Yuefeng Chen, Ranjie Duan, Yao Zhu et al.NeurIPS 2022 · 48 citations
- Towards Robust Recommendation via Decision Boundary-aware Graph Contrastive LearningJiakai Tang, Sunhao Dai, Zexu Sun, Xu Chen et al.KDD 2024 · 15 citations
- Towards Better Robustness against Common Corruptions for Unsupervised Domain AdaptationZhiqiang Gao, Kaizhu Huang, Rui Zhang, Dawei Liu et al.ICCV 2023 · 8 citations
- Training for Stable Explanation for FreeChao Chen, Chenghua Guo, Rufeng Chen, Guixiang Ma et al.NeurIPS 2024 · 7 citations
- On the Interaction of Compressibility and Adversarial RobustnessMelih Barsbey, Antônio H. Ribeiro, Umut Simsekli, Tolga BirdalICLR 2026 · 3 citations
Builds on1
Related papers
- MaxUp: Lightweight Adversarial Training With Data Augmentation Improves Neural Network TrainingChengyue Gong, Tongzheng Ren, Mao Ye, Qiang LiuCVPR 2021
- Adversarial Unlearning: Reducing Confidence Along Adversarial DirectionsAmrith Setlur, Benjamin Eysenbach, Virginia Smith, Sergey LevineNeurIPS 2022 · 26 citations
- Jacobian Adversarially Regularized Networks for RobustnessAlvin Chan, Yi Tay, Yew-Soon Ong, Jie FuICLR 2020 · 81 citations
- ε-weakened robustness of deep neural networksPei Huang, Yuting Yang, Minghao Liu, Fuqi Jia et al.ISSTA 2022 · 10 citations
- Splitting the Difference on Adversarial TrainingMatan Levi, Aryeh KontorovichUSENIX Security 2024 · 9 citations
