To be Robust or to be Fair: Towards Fairness in Adversarial Training
Han Xu, Xiaorui Liu, Yaxin Li, Anil K. Jain, Jiliang Tang
Abstract
Adversarial training algorithms have been proved to be reliable to improve machine learning models' robustness against adversarial examples. However, we find that adversarial training algorithms tend to introduce severe disparity of accuracy and robustness between different groups of data. For instance, a PGD adversarially trained ResNet18 model on CIFAR-10 has 93% clean accuracy and 67% PGD l ∞ -8 robust accuracy on the class "automobile" but only 65% and 17% on the class "cat". This phenomenon happens in balanced datasets and does not exist in naturally trained models when only using clean samples. In this work, we empirically and theoretically show that this phenomenon can happen under general adversarial training algorithms which minimize DNN models' robust errors. Motivated by these findings, we propose a Fair-Robust-Learning (FRL) framework to mitigate this unfairness problem when doing adversarial defenses. Experimental results validate the effectiveness of FRL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8f3f5a58-7917-43fc-b6b2-d093e70a245eCited by top-tier papers42
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang et al.NeurIPS 2021 · 90 citations
- Diagnosing failures of fairness transfer across distribution shift in real-world medical settingsJessica Schrouff, Natalie Harris, Sanmi Koyejo, Ibrahim M. Alabdulmohsin et al.NeurIPS 2022 · 84 citations
- Pruning has a disparate impact on model accuracyCuong Tran, Ferdinando Fioretto, Jung-Eun Kim, Rakshit NaiduNeurIPS 2022 · 64 citations
- On the Tradeoff Between Robustness and FairnessXinsong Ma, Zekai Wang, Weiwei LiuNeurIPS 2022 · 64 citations
- CalFAT: Calibrated Federated Adversarial Training with Label SkewnessChen Chen, Yuchen Liu, Xingjun Ma, Lingjuan LyuNeurIPS 2022 · 53 citations
Builds on3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 1,765 citations
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot et al.ICML 2020 · 103 citations
Related papers
- Towards Fairness-Aware Adversarial LearningYanghao Zhang, Tianle Zhang, Ronghui Mu, Xiaowei Huang et al.CVPR 2024 · 6 citations
- CFA: Class-Wise Calibrated Fair Adversarial TrainingZeming Wei, Yifei Wang, Yiwen Guo, Yisen WangCVPR 2023
- On the Alignment between Fairness and Accuracy: from the Perspective of Adversarial RobustnessJunyi Chai, Taeuk Jang, Jing Gao, Xiaoqian WangICML 2025
- Analysis and Applications of Class-wise Robustness in Adversarial TrainingQi Tian, Kun Kuang, Kelu Jiang, Fei Wu et al.KDD 2021 · 30 citations
- DAFA: Distance-Aware Fair Adversarial TrainingHyungyu Lee, Saehyung Lee, Hyemi Jang, Junsung Park et al.ICLR 2024 · 12 citations
