To be Robust or to be Fair: Towards Fairness in Adversarial Training
Han Xu, Xiaorui Liu, Yaxin Li, Anil K. Jain, Jiliang Tang
摘要
Adversarial training algorithms have been proved to be reliable to improve machine learning models' robustness against adversarial examples. However, we find that adversarial training algorithms tend to introduce severe disparity of accuracy and robustness between different groups of data. For instance, a PGD adversarially trained ResNet18 model on CIFAR-10 has 93% clean accuracy and 67% PGD l ∞ -8 robust accuracy on the class "automobile" but only 65% and 17% on the class "cat". This phenomenon happens in balanced datasets and does not exist in naturally trained models when only using clean samples. In this work, we empirically and theoretically show that this phenomenon can happen under general adversarial training algorithms which minimize DNN models' robust errors. Motivated by these findings, we propose a Fair-Robust-Learning (FRL) framework to mitigate this unfairness problem when doing adversarial defenses. Experimental results validate the effectiveness of FRL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang 等NeurIPS 2021 · 被引用 90 次
- Diagnosing failures of fairness transfer across distribution shift in real-world medical settingsJessica Schrouff, Natalie Harris, Sanmi Koyejo, Ibrahim M. Alabdulmohsin 等NeurIPS 2022 · 被引用 84 次
- Pruning has a disparate impact on model accuracyCuong Tran, Ferdinando Fioretto, Jung-Eun Kim, Rakshit NaiduNeurIPS 2022 · 被引用 64 次
- On the Tradeoff Between Robustness and FairnessXinsong Ma, Zekai Wang, Weiwei LiuNeurIPS 2022 · 被引用 64 次
- CalFAT: Calibrated Federated Adversarial Training with Label SkewnessChen Chen, Yuchen Liu, Xingjun Ma, Lingjuan LyuNeurIPS 2022 · 被引用 53 次
它引用的顶会 Paper3
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Accessorize to a Crime: Real and Stealthy Attacks on State-of-the-Art Face RecognitionMahmood Sharif, Sruti Bhagavatula, Lujo Bauer, Michael K. ReiterCCS 2016 · 被引用 1,765 次
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot 等ICML 2020 · 被引用 103 次
相关 Paper
- Towards Fairness-Aware Adversarial LearningYanghao Zhang, Tianle Zhang, Ronghui Mu, Xiaowei Huang 等CVPR 2024 · 被引用 6 次
- CFA: Class-Wise Calibrated Fair Adversarial TrainingZeming Wei, Yifei Wang, Yiwen Guo, Yisen WangCVPR 2023
- On the Alignment between Fairness and Accuracy: from the Perspective of Adversarial RobustnessJunyi Chai, Taeuk Jang, Jing Gao, Xiaoqian WangICML 2025
- Analysis and Applications of Class-wise Robustness in Adversarial TrainingQi Tian, Kun Kuang, Kelu Jiang, Fei Wu 等KDD 2021 · 被引用 30 次
- DAFA: Distance-Aware Fair Adversarial TrainingHyungyu Lee, Saehyung Lee, Hyemi Jang, Junsung Park 等ICLR 2024 · 被引用 12 次
