Improving Generalization of Adversarial Training via Robust Critical Fine-Tuning
Kaijie Zhu, Xixu Hu, Jindong Wang, Xing Xie, Ge Yang
摘要
Deep neural networks are susceptible to adversarial examples, posing a significant security risk in critical applications. Adversarial Training (AT) is a well-established technique to enhance adversarial robustness, but it often comes at the cost of decreased generalization ability. This paper proposes Robustness Critical Fine-Tuning (RiFT), a novel approach to enhance generalization without compromising adversarial robustness. The core idea of RiFT is to exploit the redundant capacity for robustness by finetuning the adversarially trained model on its non-robustcritical module. To do so, we introduce module robust criticality (MRC), a measure that evaluates the significance of a given module to model robustness under worstcase weight perturbations. Using this measure, we identify the module with the lowest MRC value as the non-robustcritical module and fine-tune its weights to obtain fine-tuned weights. Subsequently, we linearly interpolate between the adversarially trained weights and fine-tuned weights to derive the optimal fine-tuned model weights. We demonstrate the efficacy of RiFT on ResNet18, ResNet34, and WideResNet34-10 models trained on CIFAR10, CIFAR100, and Tiny-ImageNet datasets. Our experiments show that RiFT can significantly improve both generalization and outof-distribution robustness by around 1.5% while maintaining or even slightly enhancing adversarial robustness. Code is available at https://github.com/microsoft/ robustlearn .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Securely Fine-tuning Pre-trained Encoders Against Adversarial ExamplesZiqi Zhou, Minghui Li, Wei Liu, Shengshan Hu 等S&P 2024 · 被引用 23 次
- Mitigating Feature Gap for Adversarial Robustness by Feature DisentanglementNuoyan Zhou, Dawei Zhou, Decheng Liu, Nannan Wang 等AAAI 2025 · 被引用 3 次
- Sustainable Self-evolution Adversarial TrainingWenxuan Wang, Chenglei Wang, Huihui Qi, Menghao Ye 等ACM MM 2024 · 被引用 2 次
- Ciard: Cyclic Iterative Adversarial Robustness DistillationLiming Lu, Shuchao Pang, Xu Zheng, Xiang Gu 等ICCV 2025 · 被引用 1 次
- Adversarial Robust Memory-Based Continual LearnerXiaoyue Mi, Fan Tang, Zonghan Yang, Danding Wang 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 被引用 935 次
相关 Paper
- How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan 等NeurIPS 2021 · 被引用 80 次
- Boosting Adversarial Robustness with CLAT: Criticality Leveraged Adversarial TrainingBhavna Gopal, Huanrui Yang, Jingyang Zhang, Mark Horton 等ICML 2025
- Once-for-All Adversarial Training: In-Situ Tradeoff between Robustness and Accuracy for FreeHaotao Wang, Tianlong Chen, Shupeng Gui, Ting-Kuei Hu 等NeurIPS 2020 · 被引用 94 次
- Sparsity Winning Twice: Better Robust Generalization from More Efficient TrainingTianlong Chen, Zhenyu Zhang, Pengjun Wang, Santosh Balachandra 等ICLR 2022 · 被引用 54 次
- Learn2Perturb: An End-to-End Feature Perturbation Learning to Improve Adversarial RobustnessAhmadreza Jeddi, Mohammad Javad Shafiee, Michelle Karg, Christian Scharfenberger 等CVPR 2020
