CFA: Class-Wise Calibrated Fair Adversarial Training
Zeming Wei, Yifei Wang, Yiwen Guo, Yisen Wang
Abstract
Adversarial training has been widely acknowledged as the most effective method to improve the adversarial robustness against adversarial examples for Deep Neural Networks (DNNs). So far, most existing works focus on enhancing the overall model robustness, treating each class equally in both the training and testing phases. Although revealing the disparity in robustness among classes, few works try to make adversarial training fair at the class level without sacrificing overall robustness. In this paper, we are the first to theoretically and empirically investigate the preference of different classes for adversarial configurations, including perturbation margin, regularization, and weight averaging. Motivated by this, we further propose a Class-wise calibrated Fair Adversarial training framework, named CFA, which customizes specific training configurations for each class automatically. Experiments on benchmark datasets demonstrate that our proposed CFA can improve both overall robustness and fairness notably over other state-of-the-art methods. Code is available at https://github.com/PKU-ML/CFA .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Rethinking Model Ensemble in Transfer-based Adversarial AttacksHuanran Chen, Yichi Zhang, Yinpeng Dong, Xiao Yang et al.ICLR 2024 · 112 citations
- On the Duality Between Sharpness-Aware Minimization and Adversarial TrainingYihao Zhang, Hangzhou He, Jingyu Zhu, Huanran Chen et al.ICML 2024 · 29 citations
- Revisiting Adversarial Robustness Distillation from the Perspective of Robust FairnessXinli Yue, Ningping Mou, Qian Wang, Lingchen ZhaoNeurIPS 2023 · 28 citations
- Balance, Imbalance, and Rebalance: Understanding Robust Overfitting from a Minimax Game PerspectiveYifei Wang, Liangchen Li, Jiansheng Yang, Zhouchen Lin et al.NeurIPS 2023 · 26 citations
- Adversarial Examples Are Not Real FeaturesAng Li, Yifei Wang, Yiwen Guo, Yisen WangNeurIPS 2023 · 24 citations
Builds on14
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Adversarial Weight Perturbation Helps Robust GeneralizationDongxian Wu, Shu-Tao Xia, Yisen WangNeurIPS 2020 · 917 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
Related papers
- DAFA: Distance-Aware Fair Adversarial TrainingHyungyu Lee, Saehyung Lee, Hyemi Jang, Junsung Park et al.ICLR 2024 · 12 citations
- Analysis and Applications of Class-wise Robustness in Adversarial TrainingQi Tian, Kun Kuang, Kelu Jiang, Fei Wu et al.KDD 2021 · 30 citations
- Towards Fairness-Aware Adversarial LearningYanghao Zhang, Tianle Zhang, Ronghui Mu, Xiaowei Huang et al.CVPR 2024 · 6 citations
- To be Robust or to be Fair: Towards Fairness in Adversarial TrainingHan Xu, Xiaorui Liu, Yaxin Li, Anil K. Jain et al.ICML 2021 · 218 citations
- On the Tradeoff Between Robustness and FairnessXinsong Ma, Zekai Wang, Weiwei LiuNeurIPS 2022 · 64 citations
