Formulating Robustness Against Unforeseen Attacks
Sihui Dai, Saeed Mahloujifar, Prateek Mittal
摘要
Existing defenses against adversarial examples such as adversarial training typically assume that the adversary will conform to a specific or known threat model, such as p perturbations within a fixed budget. In this paper, we focus on the scenario where there is a mismatch in the threat model assumed by the defense during training, and the actual capabilities of the adversary at test time. We ask the question: if the learner trains against a specific "source" threat model, when can we expect robustness to generalize to a stronger unknown "target" threat model during test-time? Our key contribution is to formally define the problem of learning and generalization with an unforeseen adversary, which helps us reason about the increase in adversarial risk from the conventional perspective of a known adversary. Applying our framework, we derive a generalization bound which relates the generalization gap between source and target threat models to variation of the feature extractor, which measures the expected maximum difference between extracted features across a given threat model. Based on our generalization bound, we propose variation regularization (VR) which reduces variation of the feature extractor across the source threat model during training. We empirically demonstrate that using VR can lead to improved generalization to unforeseen attacks during test-time, and combining VR with perceptual adversarial training (Laidlaw et al., 2021) achieves state-of-the-art robustness on unforeseen attacks. Our code is publicly available at https://github.com/inspire-group/variation-regularization . 1. When can we expect robustness on the source threat model to generalize to the true unknown target threat model used by the adversary? 2. How can we design a learning algorithm that reduces the drop in robustness from source threat model to target threat model? 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- OODRobustBench: a Benchmark and Large-Scale Analysis of Adversarial Robustness under Distribution ShiftLin Li, Yifei Wang, Chawin Sitawarin, Michael W. SpratlingICML 2024 · 被引用 13 次
- MultiRobustBench: Benchmarking Robustness Against Multiple AttacksSihui Dai, Saeed Mahloujifar, Chong Xiang, Vikash Sehwag 等ICML 2023 · 被引用 11 次
- Adapting to Evolving Adversaries with Regularized Continual Robust TrainingSihui Dai, Christian Cianfarani, Vikash Sehwag, Prateek Mittal 等ICML 2025
它引用的顶会 Paper12
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 被引用 917 次
- Randomized Smoothing of All Shapes and SizesGreg Yang, Tony Duan, J. Edward Hu, Hadi Salman 等ICML 2020 · 被引用 237 次
- Perceptual Adversarial Robustness: Defense Against Unseen Threat ModelsCassidy Laidlaw, Sahil Singla, Soheil FeiziICLR 2021 · 被引用 217 次
相关 Paper
- Towards Better Robust Generalization with Shift Consistency RegularizationShufei Zhang, Zhuang Qian, Kaizhu Huang, Qiufeng Wang 等ICML 2021 · 被引用 18 次
- Robustness and Generalization via Generative Adversarial TrainingOmid Poursaeed, Tianxing Jiang, Harry Yang, Serge J. Belongie 等ICCV 2021 · 被引用 35 次
- Certifying Better Robust Generalization for Unsupervised Domain AdaptationZhiqiang Gao, Shufei Zhang, Kaizhu Huang, Qiufeng Wang 等ACM MM 2022 · 被引用 4 次
- Learning Unforeseen Robustness from Out-of-distribution Data Using Equivariant Domain TranslatorSicheng Zhu, Bang An, Furong Huang, Sanghyun HongICML 2023 · 被引用 4 次
- Improving Adversarial Robustness by Putting More Regularizations on Less Robust SamplesDongyoon Yang, Insung Kong, Yongdai KimICML 2023 · 被引用 15 次
