Adapting to Evolving Adversaries with Regularized Continual Robust Training
Sihui Dai, Christian Cianfarani, Vikash Sehwag, Prateek Mittal, Arjun Nitin Bhagoji
Abstract
Robust training methods typically defend against specific attack types, such as ℓ p attacks with fixed budgets, and rarely account for the fact that defenders may encounter new attacks over time. A natural solution is to adapt the defended model to new adversaries as they arise via fine-tuning, a method which we call continual robust training (CRT). However, when implemented naively, fine-tuning on new attacks degrades robustness on previous attacks. This raises the question: how can we improve the initial training and fine-tuning of the model to simultaneously achieve robustness against previous and new attacks? We present theoretical results which show that the gap in a model's robustness against different attacks is bounded by how far each attack perturbs a sample in the model's logit space, suggesting that regularizing with respect to this logit space distance can help maintain robustness against previous attacks. Extensive experiments on 3 datasets (CIFAR-10, CIFAR-100, and ImageNette) and over 100 attack combinations demonstrate that the proposed regularization improves robust accuracy with little overhead in training time. Our findings and open-source code 1 lay the groundwork for the deployment of models robust to evolving attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2d6c53d-0fb3-4626-8a74-aad8fdc6d987Builds on23
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
- Understanding Self-Training for Gradual Domain AdaptationAnanya Kumar, Tengyu Ma, Percy LiangICML 2020 · 266 citations
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 263 citations
Related papers
- Exploring The Forgetting in Adversarial Training: A Novel Method for Enhancing RobustnessXianglu Wang, Hu DingICLR 2025
- Seasoning Model Soups for Robustness to Adversarial and Natural Distribution ShiftsFrancesco Croce, Sylvestre-Alvise Rebuffi, Evan Shelhamer, Sven GowalCVPR 2023
- RAMP: Boosting Adversarial Robustness Against Multiple lp Perturbations for Universal RobustnessEnyi Jiang, Gagandeep SinghNeurIPS 2024
- MultiRobustBench: Benchmarking Robustness Against Multiple AttacksSihui Dai, Saeed Mahloujifar, Chong Xiang, Vikash Sehwag et al.ICML 2023 · 11 citations
- Towards Better Robust Generalization with Shift Consistency RegularizationShufei Zhang, Zhuang Qian, Kaizhu Huang, Qiufeng Wang et al.ICML 2021 · 18 citations
