Low Curvature Activations Reduce Overfitting in Adversarial Training
Vasu Singla, Sahil Singla, Soheil Feizi, David Jacobs
Abstract
Adversarial training is one of the most effective defenses against adversarial attacks. Previous works suggest that overfitting is a dominant phenomenon in adversarial training leading to a large generalization gap between test and train accuracy in neural networks. In this work, we show that the observed generalization gap is closely related to the choice of the activation function. In particular, we show that using activation functions with low (exact or approximate) curvature values has a regularization effect that significantly reduces both the standard and robust generalization gaps in adversarial training. We observe this effect for both differentiable/smooth activations such as SiLU as well as non-differentiable/non-smooth activations such as LeakyReLU. In the latter case, the "approximate" curvature of the activation is low. Finally, we show that for activation functions with low curvature, the double descent phenomenon for adversarially trained models does not occur.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a414e58e-4a82-4fbd-9c5e-bad5eee296adCited by top-tier papers11
- Robustness and Accuracy Could Be Reconcilable by (Proper) DefinitionTianyu Pang, Min Lin, Xiao Yang, Jun Zhu et al.ICML 2022 · 168 citations
- Relating Adversarially Robust Generalization to Flat MinimaDavid Stutz, Matthias Hein, Bernt SchieleICCV 2021 · 80 citations
- Sparsity Winning Twice: Better Robust Generalization from More Efficient TrainingTianlong Chen, Zhenyu Zhang, Pengjun Wang, Santosh Balachandra et al.ICLR 2022 · 54 citations
- Label Noise in Adversarial Training: A Novel Perspective to Study Robust OverfittingChengyu Dong, Liyuan Liu, Jingbo ShangNeurIPS 2022 · 36 citations
- Efficient local linearity regularization to overcome catastrophic overfittingElías Abad-Rocamora, Fanghui Liu, Grigorios Chrysos, Pablo M. Olmos et al.ICLR 2024 · 9 citations
Builds on23
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
Related papers
- Benign Overfitting in Adversarial Training of Neural NetworksYunjuan Wang, Kaibo Zhang, Raman AroraICML 2024 · 3 citations
- Efficient Training of Low-Curvature Neural NetworksSuraj Srinivas, Kyle Matoba, Himabindu Lakkaraju, François FleuretNeurIPS 2022 · 23 citations
- Over-parameterized Adversarial Training: An Analysis Overcoming the Curse of DimensionalityYi Zhang, Orestis Plevrakis, Simon S. Du, Xingguo Li et al.NeurIPS 2020 · 56 citations
- Effect of Activation Functions on the Training of Overparametrized Neural NetsAbhishek Panigrahi, Abhishek Shetty, Navin GoyalICLR 2020 · 24 citations
- Enhancing Adversarial Defense by k-Winners-Take-AllChang Xiao, Peilin Zhong, Changxi ZhengICLR 2020 · 114 citations
