On the Alignment between Fairness and Accuracy: from the Perspective of Adversarial Robustness
Junyi Chai, Taeuk Jang, Jing Gao, Xiaoqian Wang
Abstract
While numerous work has been proposed to address fairness in machine learning, existing methods do not guarantee fair predictions under imperceptible feature perturbation, and a seemingly fair model can suffer from large group-wise disparities under such perturbation. Moreover, while adversarial training has been shown to be reliable in improving a model's robustness to defend against adversarial feature perturbation that deteriorates accuracy, it has not been properly studied in the context of adversarial perturbation against fairness. To tackle these challenges, in this paper, we study the problem of adversarial attack and adversarial robustness w.r.t. two terms: fairness and accuracy. From the adversarial attack perspective, we propose a unified structure for adversarial attacks against fairness which brings together common notions in group fairness, and we theoretically prove the equivalence of adversarial attacks against different fairness notions. Further, we derive the connections between adversarial attacks against fairness and those against accuracy. From the adversarial robustness perspective, we theoretically align robustness to adversarial attacks against fairness and accuracy, where robustness w.r.t. one term enhances robustness w.r.t. the other term. Our study suggests a novel way to unify adversarial training w.r.t. fairness and accuracy, and experiments show our proposed method achieves better robustness w.r.t. both terms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c5ec549a-092c-4fdd-bd1b-63429b1e0c81Builds on23
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey et al.ICLR 2020 · 829 citations
- To be Robust or to be Fair: Towards Fairness in Adversarial TrainingHan Xu, Xiaorui Liu, Yaxin Li, Anil K. Jain et al.ICML 2021 · 218 citations
- Learnable Boundary Guided Adversarial TrainingJiequan Cui, Shu Liu, Liwei Wang, Jiaya JiaICCV 2021 · 152 citations
- Two Simple Ways to Learn Individual Fairness Metrics from DataDebarghya Mukherjee, Mikhail Yurochkin, Moulinath Banerjee, Yuekai SunICML 2020 · 109 citations
Related papers
- On the Tradeoff Between Robustness and FairnessXinsong Ma, Zekai Wang, Weiwei LiuNeurIPS 2022 · 64 citations
- DAFA: Distance-Aware Fair Adversarial TrainingHyungyu Lee, Saehyung Lee, Hyemi Jang, Junsung Park et al.ICLR 2024 · 12 citations
- CFA: Class-Wise Calibrated Fair Adversarial TrainingZeming Wei, Yifei Wang, Yiwen Guo, Yisen WangCVPR 2023
- Can Private Machine Learning Be Fair?Joseph Rance, Filip SvobodaAAAI 2025 · 1 citation
- Towards Fairness-Aware Adversarial LearningYanghao Zhang, Tianle Zhang, Ronghui Mu, Xiaowei Huang et al.CVPR 2024 · 6 citations
