Towards Poisoning Fair Representations
Tianci Liu, Haoyu Wang, Feijie Wu, Hengtong Zhang, Pan Li, Lu Su, Jing Gao
摘要
Fair machine learning seeks to mitigate model prediction bias against certain demographic subgroups such as elder and female. Recently, fair representation learning (FRL) trained by deep neural networks has demonstrated superior performance, whereby representations containing no demographic information are inferred from the data and then used as the input to classification or other downstream tasks. Despite the development of FRL methods, their vulnerability under data poisoning attack, a popular protocol to benchmark model robustness under adversarial scenarios, is under-explored. Data poisoning attacks have been developed for classical fair machine learning methods which incorporate fairness constraints into shallow-model classifiers. Nonetheless, these attacks fall short in FRL due to notably different fairness goals and model architectures. This work proposes the first data poisoning framework attacking FRL. We induce the model to output unfair representations that contain as much demographic information as possible by injecting carefully crafted poisoning samples into the training data. This attack entails a prohibitive bilevel optimization, wherefore an effective approximated solution is proposed. A theoretical analysis on the needed number of poisoning samples is derived and sheds light on defending against the attack. Experiments on benchmark fairness datasets and state-of-the-art fair representation learning models demonstrate the superiority of our attack.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Witches' Brew: Industrial Scale Data Poisoning via Gradient MatchingJonas Geiping, Liam H. Fowl, W. Ronny Huang, Wojciech Czaja 等ICLR 2021 · 被引用 268 次
- MetaPoison: Practical General-purpose Clean-label Data PoisoningW. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor 等NeurIPS 2020 · 被引用 242 次
- Conditional Learning of Fair RepresentationsHan Zhao, Amanda Coston, Tameem Adel, Geoffrey J. GordonICLR 2020 · 被引用 127 次
- Controllable Guarantees for Fair Outcomes via Contrastive Information EstimationUmang Gupta, Aaron M. Ferber, Bistra Dilkina, Greg Ver SteegAAAI 2021 · 被引用 78 次
- Exacerbating Algorithmic Bias through Fairness AttacksNinareh Mehrabi, Muhammad Naveed, Fred Morstatter, Aram GalstyanAAAI 2021 · 被引用 76 次
相关 Paper
- Fair Representation Learning with Controllable High Confidence Guarantees via Adversarial InferenceYuhong Luo, Austin Hoag, Xintong Wang, Philip S. Thomas 等NeurIPS 2025
- Sustaining Fairness via Incremental LearningSomnath Basu Roy Chowdhury, Snigdha ChaturvediAAAI 2023 · 被引用 6 次
- Learning for Counterfactual Fairness from Observational DataJing Ma, Ruocheng Guo, Aidong Zhang, Jundong LiKDD 2023 · 被引用 9 次
- First-Order Efficient General-Purpose Clean-Label Data PoisoningTianhang Zheng, Baochun LiINFOCOM 2021 · 被引用 8 次
- Fairness ReprogrammingGuanhua Zhang, Yihua Zhang, Yang Zhang, Wenqi Fan 等NeurIPS 2022 · 被引用 46 次
