Improving Generalization of Adversarial Training via Robust Critical Fine-Tuning
Kaijie Zhu, Xixu Hu, Jindong Wang, Xing Xie, Ge Yang
Abstract
Deep neural networks are susceptible to adversarial examples, posing a significant security risk in critical applications. Adversarial Training (AT) is a well-established technique to enhance adversarial robustness, but it often comes at the cost of decreased generalization ability. This paper proposes Robustness Critical Fine-Tuning (RiFT), a novel approach to enhance generalization without compromising adversarial robustness. The core idea of RiFT is to exploit the redundant capacity for robustness by finetuning the adversarially trained model on its non-robustcritical module. To do so, we introduce module robust criticality (MRC), a measure that evaluates the significance of a given module to model robustness under worstcase weight perturbations. Using this measure, we identify the module with the lowest MRC value as the non-robustcritical module and fine-tune its weights to obtain fine-tuned weights. Subsequently, we linearly interpolate between the adversarially trained weights and fine-tuned weights to derive the optimal fine-tuned model weights. We demonstrate the efficacy of RiFT on ResNet18, ResNet34, and WideResNet34-10 models trained on CIFAR10, CIFAR100, and Tiny-ImageNet datasets. Our experiments show that RiFT can significantly improve both generalization and outof-distribution robustness by around 1.5% while maintaining or even slightly enhancing adversarial robustness. Code is available at https://github.com/microsoft/ robustlearn .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81a1d5c4-9593-4e05-9e35-b58750bd15b0Cited by top-tier papers14
- Securely Fine-tuning Pre-trained Encoders Against Adversarial ExamplesZiqi Zhou, Minghui Li, Wei Liu, Shengshan Hu et al.S&P 2024 · 23 citations
- Mitigating Feature Gap for Adversarial Robustness by Feature DisentanglementNuoyan Zhou, Dawei Zhou, Decheng Liu, Nannan Wang et al.AAAI 2025 · 3 citations
- Sustainable Self-evolution Adversarial TrainingWenxuan Wang, Chenglei Wang, Huihui Qi, Menghao Ye et al.ACM MM 2024 · 2 citations
- Ciard: Cyclic Iterative Adversarial Robustness DistillationLiming Lu, Shuchao Pang, Xu Zheng, Xiang Gu et al.ICCV 2025 · 1 citation
- Adversarial Robust Memory-Based Continual LearnerXiaoyue Mi, Fan Tang, Zonghan Yang, Danding Wang et al.ICCV 2025 · 1 citation
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
Related papers
- How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan et al.NeurIPS 2021 · 80 citations
- Boosting Adversarial Robustness with CLAT: Criticality Leveraged Adversarial TrainingBhavna Gopal, Huanrui Yang, Jingyang Zhang, Mark Horton et al.ICML 2025
- Once-for-All Adversarial Training: In-Situ Tradeoff between Robustness and Accuracy for FreeHaotao Wang, Tianlong Chen, Shupeng Gui, Ting-Kuei Hu et al.NeurIPS 2020 · 94 citations
- Sparsity Winning Twice: Better Robust Generalization from More Efficient TrainingTianlong Chen, Zhenyu Zhang, Pengjun Wang, Santosh Balachandra et al.ICLR 2022 · 54 citations
- Learn2Perturb: An End-to-End Feature Perturbation Learning to Improve Adversarial RobustnessAhmadreza Jeddi, Mohammad Javad Shafiee, Michelle Karg, Christian Scharfenberger et al.CVPR 2020
