Towards Poisoning Fair Representations
Tianci Liu, Haoyu Wang, Feijie Wu, Hengtong Zhang, Pan Li, Lu Su, Jing Gao
Abstract
Fair machine learning seeks to mitigate model prediction bias against certain demographic subgroups such as elder and female. Recently, fair representation learning (FRL) trained by deep neural networks has demonstrated superior performance, whereby representations containing no demographic information are inferred from the data and then used as the input to classification or other downstream tasks. Despite the development of FRL methods, their vulnerability under data poisoning attack, a popular protocol to benchmark model robustness under adversarial scenarios, is under-explored. Data poisoning attacks have been developed for classical fair machine learning methods which incorporate fairness constraints into shallow-model classifiers. Nonetheless, these attacks fall short in FRL due to notably different fairness goals and model architectures. This work proposes the first data poisoning framework attacking FRL. We induce the model to output unfair representations that contain as much demographic information as possible by injecting carefully crafted poisoning samples into the training data. This attack entails a prohibitive bilevel optimization, wherefore an effective approximated solution is proposed. A theoretical analysis on the needed number of poisoning samples is derived and sheds light on defending against the attack. Experiments on benchmark fairness datasets and state-of-the-art fair representation learning models demonstrate the superiority of our attack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5afb36f9-d74d-424f-97b5-7f0caae4c16bCited by top-tier papers1
Ask how each one uses itBuilds on7
- Witches' Brew: Industrial Scale Data Poisoning via Gradient MatchingJonas Geiping, Liam H. Fowl, W. Ronny Huang, Wojciech Czaja et al.ICLR 2021 · 268 citations
- MetaPoison: Practical General-purpose Clean-label Data PoisoningW. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor et al.NeurIPS 2020 · 242 citations
- Conditional Learning of Fair RepresentationsHan Zhao, Amanda Coston, Tameem Adel, Geoffrey J. GordonICLR 2020 · 127 citations
- Controllable Guarantees for Fair Outcomes via Contrastive Information EstimationUmang Gupta, Aaron M. Ferber, Bistra Dilkina, Greg Ver SteegAAAI 2021 · 78 citations
- Exacerbating Algorithmic Bias through Fairness AttacksNinareh Mehrabi, Muhammad Naveed, Fred Morstatter, Aram GalstyanAAAI 2021 · 76 citations
Related papers
- Fair Representation Learning with Controllable High Confidence Guarantees via Adversarial InferenceYuhong Luo, Austin Hoag, Xintong Wang, Philip S. Thomas et al.NeurIPS 2025
- Sustaining Fairness via Incremental LearningSomnath Basu Roy Chowdhury, Snigdha ChaturvediAAAI 2023 · 6 citations
- Learning for Counterfactual Fairness from Observational DataJing Ma, Ruocheng Guo, Aidong Zhang, Jundong LiKDD 2023 · 9 citations
- First-Order Efficient General-Purpose Clean-Label Data PoisoningTianhang Zheng, Baochun LiINFOCOM 2021 · 8 citations
- Fairness ReprogrammingGuanhua Zhang, Yihua Zhang, Yang Zhang, Wenqi Fan et al.NeurIPS 2022 · 46 citations
