On the Tradeoff Between Robustness and Fairness
Xinsong Ma, Zekai Wang, Weiwei Liu
Abstract
Interestingly, recent experimental results [2, 26] have identified a robust fairness phenomenon in adversarial training (AT), namely that a robust model well-trained by AT exhibits a remarkable disparity of standard accuracy and robust accuracy among different classes compared with natural training. However, the effect of different perturbation radii in AT on robust fairness has not been studied, and one natural question is raised: does a tradeoff exist between average robustness and robust fairness? Our extensive experimental results provide an affirmative answer to this question: with an increasing perturbation radius, stronger AT will lead to a larger class-wise disparity of robust accuracy. Theoretically, we analyze the class-wise performance of adversarially trained linear models with mixture Gaussian distribution. Our theoretical results support our observations. Moreover, our theory shows that adversarial training easily leads to more serious robust fairness issue than natural training. Motivated by theoretical results, we propose a fairly adversarial training (FAT) method to mitigate the tradeoff between average robustness and robust fairness. Experimental results validate the effectiveness of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e30c1905-89ab-4441-87f2-50286aae78c2Cited by top-tier papers31
- Better Diffusion Models Further Improve Adversarial TrainingZekai Wang, Tianyu Pang, Chao Du, Min Lin et al.ICML 2023 · 300 citations
- Adversarial Self-Training Improves Robustness and Generalization for Gradual Domain AdaptationLianghe Shi, Weiwei LiuNeurIPS 2023 · 34 citations
- Revisiting Adversarial Robustness Distillation from the Perspective of Robust FairnessXinli Yue, Ningping Mou, Qian Wang, Lingchen ZhaoNeurIPS 2023 · 28 citations
- A Theory of Transfer-Based Black-Box Attacks: Explanation and ImplicationsYanbo Chen, Weiwei LiuNeurIPS 2023 · 22 citations
- H-Consistency Guarantees for RegressionAnqi Mao, Mehryar Mohri, Yutao ZhongICML 2024 · 18 citations
Builds on5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- To be Robust or to be Fair: Towards Fairness in Adversarial TrainingHan Xu, Xiaorui Liu, Yaxin Li, Anil K. Jain et al.ICML 2021 · 218 citations
- Boosting Adversarial Training with Hypersphere EmbeddingTianyu Pang, Xiao Yang, Yinpeng Dong, Taufik Xu et al.NeurIPS 2020 · 170 citations
- Robustness Verification for Contrastive LearningZekai Wang, Weiwei LiuICML 2022 · 17 citations
Related papers
- DAFA: Distance-Aware Fair Adversarial TrainingHyungyu Lee, Saehyung Lee, Hyemi Jang, Junsung Park et al.ICLR 2024 · 12 citations
- Understanding the Impact of Adversarial Robustness on Accuracy DisparityYuzheng Hu, Fan Wu, Hongyang Zhang, Han ZhaoICML 2023 · 11 citations
- CFA: Class-Wise Calibrated Fair Adversarial TrainingZeming Wei, Yifei Wang, Yiwen Guo, Yisen WangCVPR 2023
- On the Alignment between Fairness and Accuracy: from the Perspective of Adversarial RobustnessJunyi Chai, Taeuk Jang, Jing Gao, Xiaoqian WangICML 2025
- Towards Fairness-Aware Adversarial LearningYanghao Zhang, Tianle Zhang, Ronghui Mu, Xiaowei Huang et al.CVPR 2024 · 6 citations
