Are Adversarial Examples Created Equal? A Learnable Weighted Minimax Risk for Robustness under Non-uniform Attacks
Huimin Zeng, Chen Zhu, Tom Goldstein, Furong Huang
Abstract
Adversarial Training is proved to be an efficient method to defend against adversarial examples, being one of the few defenses that withstand strong attacks. However, traditional defense mechanisms assume a uniform attack over the examples according to the underlying data distribution, which is apparently unrealistic as the attacker could choose to focus on more vulnerable examples. We present a weighted minimax risk optimization that defends against non-uniform attacks, achieving robustness against adversarial examples under perturbed test data distributions. Our modified risk considers importance weights of different adversarial examples and focuses adaptively on harder examples that are wrongly classified or at higher risk of being classified incorrectly. The designed risk allows the training process to learn a strong defense through optimizing the importance weights. The experiments show that our model significantly improves state-of-the-art adversarial accuracy under non-uniform attacks without a significant drop under uniform attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Revisiting Contrastive Learning through the Lens of Neighborhood Component Analysis: an Integrated FrameworkChing-Yun Ko, Jeet Mohapatra, Sijia Liu, Pin-Yu Chen et al.ICML 2022 · 15 citations
- Exploring and Exploiting Decision Boundary Dynamics for Adversarial RobustnessYuancheng Xu, Yanchao Sun, Micah Goldblum, Tom Goldstein et al.ICLR 2023 · 10 citations
- Memorization Weights for Instance Reweighting in Adversarial TrainingJianfu Zhang, Yan Hong, Qibin ZhaoAAAI 2023 · 4 citations
- Doubly Robust Instance-Reweighted Adversarial TrainingDaouda Sow, Sen Lin, Zhangyang Wang, Yingbin LiangICLR 2024 · 2 citations
- Dynamic Loss-Based Sample Reweighting for Improved Large Language Model PretrainingDaouda Sow, Herbert Woisetschläger, Saikiran Bulusu, Shiqiang Wang et al.ICLR 2025
Builds on5
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- MagNet: A Two-Pronged Defense against Adversarial ExamplesDongyu Meng, Hao ChenCCS 2017 · 1,295 citations
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu et al.NeurIPS 2020 · 561 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
- SMART: Robust and Efficient Fine-Tuning for Pre-trained Natural Language Models through Principled Regularized OptimizationHaoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu et al.ACL 2020 · 148 citations
Related papers
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang et al.NeurIPS 2021 · 90 citations
- Adversarial Attack Generation Empowered by Min-Max OptimizationJingkang Wang, Tianyun Zhang, Sijia Liu, Pin-Yu Chen et al.NeurIPS 2021 · 49 citations
- Class-Aware Robust Adversarial Training for Object DetectionPin-Chun Chen, Bo-Han Kung, Jun-Cheng ChenCVPR 2021
- Advancing Example Exploitation Can Alleviate Critical Challenges in Adversarial TrainingYao Ge, Yun Li, Keji Han, Junyi Zhu et al.ICCV 2023 · 6 citations
