Exacerbating Algorithmic Bias through Fairness Attacks
Ninareh Mehrabi, Muhammad Naveed, Fred Morstatter, Aram Galstyan
Abstract
Algorithmic fairness has attracted significant attention in recent years, with many quantitative measures suggested for characterizing the fairness of different machine learning algorithms. Despite this interest, the robustness of those fairness measures with respect to an intentional adversarial attack has not been properly addressed. Indeed, most adversarial machine learning has focused on the impact of malicious attacks on the accuracy of the system, without any regard to the system's fairness. We propose new types of data poisoning attacks where an adversary intentionally targets the fairness of a system. Specifically, we propose two families of attacks that target fairness measures. In the anchoring attack, we skew the decision boundary by placing poisoned points near specific target points to bias the outcome. In the influence attack on fairness, we aim to maximize the covariance between the sensitive attributes and the decision outcome and affect the fairness of the model. We conduct extensive experiments that indicate the effectiveness of our proposed attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8675949b-806f-4a42-a038-08a1ebf8e7bbCited by top-tier papers16
- Interpretable Data-Based Explanations for Fairness DebuggingRomila Pradhan, Jiongli Zhu, Boris Glavic, Babak SalimiSIGMOD 2022 · 53 citations
- Are My Deep Learning Systems Fair? An Empirical Study of Fixed-Seed TrainingShangshu Qian, Hung Viet Pham, Thibaud Lutellier, Zeou Hu et al.NeurIPS 2021 · 49 citations
- "What Data Benefits My Classifier?" Enhancing Model Performance and Interpretability through Influence-Based Data SelectionAnshuman Chhabra, Peizhao Li, Prasant Mohapatra, Hongfu LiuICLR 2024 · 32 citations
- Improving Fair Training under Correlation ShiftsYuji Roh, Kangwook Lee, Steven Euijong Whang, Changho SuhICML 2023 · 22 citations
- Deceptive Fairness Attacks on Graphs via Meta LearningJian Kang, Yinglong Xia, Ross Maciejewski, Jiebo Luo et al.ICLR 2024 · 9 citations
Related papers
- Towards Poisoning Fair RepresentationsTianci Liu, Haoyu Wang, Feijie Wu, Hengtong Zhang et al.ICLR 2024 · 3 citations
- Model-Targeted Poisoning Attacks with Provable ConvergenceFnu Suya, Saeed Mahloujifar, Anshuman Suri, David Evans et al.ICML 2021 · 52 citations
- Attack To Defend: Exploiting Adversarial Attacks for Detecting Poisoned ModelsSamar Fares, Karthik NandakumarCVPR 2024
- A Separation Result Between Data-oblivious and Data-aware Poisoning AttacksSamuel Deng, Sanjam Garg, Somesh Jha, Saeed Mahloujifar et al.NeurIPS 2021 · 3 citations
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
