Removing Batch Normalization Boosts Adversarial Training
Haotao Wang, Aston Zhang, Shuai Zheng, Xingjian Shi, Mu Li, Zhangyang Wang
Abstract
Adversarial training (AT) defends deep neural networks against adversarial attacks. One challenge that limits its practical application is the performance degradation on clean samples. A major bottleneck identified by previous works is the widely used batch normalization (BN), which struggles to model the different statistics of clean and adversarial training samples in AT. Although the dominant approach is to extend BN to capture this mixture of distribution, we propose to completely eliminate this bottleneck by removing all BN layers in AT. Our normalizer-free robust training (NoFrost) method extends recent advances in normalizer-free networks to AT for its unexplored advantage on handling the mixture distribution challenge. We show that NoFrost achieves adversarial robustness with only a minor sacrifice on clean sample accuracy. On ImageNet with ResNet50, NoFrost achieves clean accuracy, which drops merely from standard training. In contrast, BN-based AT obtains clean accuracy, suffering a significant drop from standard training. In addition, NoFrost achieves a adversarial robustness against PGD attack, which improves the robustness in BN-based AT. We observe better model smoothness and larger decision margins from NoFrost, which make the models less sensitive to input perturbations and thus more robust. Moreover, when incorporating more data augmentations into NoFrost, it achieves comprehensive robustness against multiple distribution shifts. Code and pre-trained models are public at https://github.com/amazon-research/normalizer-free-robust-training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext acd5c8e1-d66c-4e25-97ce-d9acc4e38541Cited by top-tier papers14
- Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity ModelingHaotao Wang, Ziyu Jiang, Yuning You, Yan Han et al.NeurIPS 2023 · 104 citations
- SNN-RAT: Robustness-enhanced Spiking Neural Network through Regularized Adversarial TrainingJianhao Ding, Tong Bu, Zhaofei Yu, Tiejun Huang et al.NeurIPS 2022 · 70 citations
- Shared Adversarial Unlearning: Backdoor Mitigation by Unlearning Shared Adversarial ExamplesShaokui Wei, Mingda Zhang, Hongyuan Zha, Baoyuan WuNeurIPS 2023 · 69 citations
- MAP: Towards Balanced Generalization of IID and OOD through Model-Agnostic AdaptersMin Zhang, Junkun Yuan, Yue He, Wenbin Li et al.ICCV 2023 · 21 citations
- Phase-aware Adversarial Defense for Improving Adversarial RobustnessDawei Zhou, Nannan Wang, Heng Yang, Xinbo Gao et al.ICML 2023 · 14 citations
Builds on25
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- Random Erasing Data AugmentationZhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li et al.AAAI 2020 · 4,134 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
Related papers
- Intriguing Properties of Adversarial Training at ScaleCihang Xie, Alan L. YuilleICLR 2020 · 66 citations
- Encoding Robustness to Image Style via Adversarial Feature PerturbationsManli Shu, Zuxuan Wu, Micah Goldblum, Tom GoldsteinNeurIPS 2021 · 23 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Delving into the Estimation Shift of Batch Normalization in a NetworkLei Huang, Yi Zhou, Tian Wang, Jie Luo et al.CVPR 2022 · 25 citations
- Adversarial Finetuning with Latent Representation Constraint to Mitigate Accuracy-Robustness TradeoffSatoshi Suzuki, Shin'ya Yamaguchi, Shoichiro Takeda, Sekitoshi Kanai et al.ICCV 2023 · 8 citations
