Fair Classification with Adversarial Perturbations
L. Elisa Celis, Anay Mehrotra, Nisheeth K. Vishnoi
Abstract
We study fair classification in the presence of an omniscient adversary that, given an , is allowed to choose an arbitrary -fraction of the training samples and arbitrarily perturb their protected attributes. The motivation comes from settings in which protected attributes can be incorrect due to strategic misreporting, malicious actors, or errors in imputation; and prior approaches that make stochastic or independence assumptions on errors may not satisfy their guarantees in this adversarial setting. Our main contribution is an optimization framework to learn fair classifiers in this adversarial setting that comes with provable guarantees on accuracy and fairness. Our framework works with multiple and non-binary protected attributes, is designed for the large class of linear-fractional fairness metrics, and can also handle perturbations besides protected attributes. We prove near-tightness of our framework's guarantees for natural hypothesis classes: no algorithm can have significantly better accuracy and any algorithm with better fairness must have lower accuracy. Empirically, we evaluate the classifiers produced by our framework for statistical rate on real-world and synthetic datasets for a family of adversaries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Fairness without Demographics through Knowledge DistillationJunyi Chai, Taeuk Jang, Xiaoqian WangNeurIPS 2022 · 57 citations
- Fair Ranking with Noisy Protected AttributesAnay Mehrotra, Nisheeth K. VishnoiNeurIPS 2022 · 24 citations
- Chasing Fairness Under Distribution Shift: A Model Weight Perturbation ApproachZhimeng Stephen Jiang, Xiaotian Han, Hongye Jin, Guanchu Wang et al.NeurIPS 2023 · 22 citations
- Input-agnostic Certified Group Fairness via Gaussian Parameter SmoothingJiayin Jin, Zeru Zhang, Yang Zhou, Lingfei WuICML 2022 · 18 citations
- Fair Classification with Partial Feedback: An Exploration-Based Data Collection ApproachVijay Keswani, Anay Mehrotra, L. Elisa CelisICML 2024 · 3 citations
Builds on2
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee et al.NeurIPS 2020 · 406 citations
- Robust Optimization for Fairness with Noisy Protected GroupsSerena Lutong Wang, Wenshuo Guo, Harikrishna Narasimhan, Andrew Cotter et al.NeurIPS 2020 · 134 citations
Related papers
- Fair Classification with Noisy Protected Attributes: A Framework with Provable GuaranteesL. Elisa Celis, Lingxiao Huang, Vijay Keswani, Nisheeth K. VishnoiICML 2021 · 67 citations
- Ensuring Fairness Beyond the Training DataDebmalya Mandal, Samuel Deng, Suman Jana, Jeannette M. Wing et al.NeurIPS 2020 · 68 citations
- Demystifying the Optimal Fair Classifier in Multi-Class ClassificationLi Zhang, Yuyuan Li, XiaoHua Feng, Jiaming Zhang et al.ICML 2026
- Individual Fairness In Strategic ClassificationZhiqun Zuo, Mohammad Mahdi KhaliliNeurIPS 2025
- Adapting Fairness Interventions to Missing ValuesRaymond Feng, Flávio P. Calmon, Hao WangNeurIPS 2023 · 20 citations
