Fairness-aware Adversarial Perturbation Towards Bias Mitigation for Deployed Deep Models
Zhibo Wang, Xiaowei Dong, Henry Xue, Zhifei Zhang, Weifeng Chiu, Tao Wei, Kui Ren
Abstract
Prioritizing fairness is of central importance in artificial intelligence (AI) systems, especially for those societal applications, e.g., hiring systems should recommend applicants equally from different demographic groups, and risk assessment systems must eliminate racism in criminal justice. Existing efforts towards the ethical development of AI systems have leveraged data science to mitigate biases in the training set or introduced fairness principles into the training process. For a deployed AI system, however, it may not allow for retraining or tuning in practice. By contrast, we propose a more flexible approach, i.e., fairness-aware adversarial perturbation (FAAP), which learns to perturb input data to blind deployed models on fairness-related features, e.g., gender and ethnicity. The key advantage is that FAAP does not modify deployed models in terms of param-eters and structures. To achieve this, we design a discriminator to distinguish fairness-related attributes based on latent representations from deployed models. Meanwhile, a perturbation generator is trained against the discriminator, such that no fairness-related features could be extracted from perturbed inputs. Exhaustive experimental evaluation demonstrates the effectiveness and superior performance of the proposed FAAP. In addition, FAAP is validated on real-world commercial deployments (inaccessible to model pa-rameters), which shows the transferability of FAAP, foreseeing the potential of black-box adaptation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- FairSeg: A Large-Scale Medical Image Segmentation Dataset for Fairness Learning Using Segment Anything Model with Fair Error-Bound ScalingYu Tian, Min Shi, Yan Luo, Ava Kouhana et al.ICLR 2024 · 9 citations
- CAD-VAE: Leveraging Correlation-Aware Latents for Comprehensive Fair DisentanglementChenrui Ma, Xi Xiao, Tianyang Wang, Xiao Wang et al.AAAI 2026 · 9 citations
- Mitigating Biases in Blackbox Feature Extractors for Image Classification TasksAbhipsa Basu, Saswat Subhajyoti Mallick, R. Venkatesh BabuNeurIPS 2024 · 6 citations
- Fair and Accurate Decision Making through Group-Aware LearningRamtin Hosseini, Li Zhang, Bhanu Garg, Pengtao XieICML 2023 · 6 citations
- Distributionally Generative Augmentation for Fair Facial Attribute ClassificationFengda Zhang, Qianpei He, Kun Kuang, Jiashuo Liu et al.CVPR 2024 · 4 citations
Builds on5
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- FR-Train: A Mutual Information-Based Approach to Fair and Robust TrainingYuji Roh, Kangwook Lee, Steven Whang, Changho SuhICML 2020 · 90 citations
- Towards Accuracy-Fairness Paradox: Adversarial Example-based Data Augmentation for Visual DebiasingYi Zhang, Jitao SangACM MM 2020 · 32 citations
- Towards Universal Representation Learning for Deep Face RecognitionYichun Shi, Xiang Yu, Kihyuk Sohn, Manmohan Chandraker et al.CVPR 2020
- Fair Attribute Classification Through Latent Space De-BiasingVikram V. Ramaswamy, Sunnie S. Y. Kim, Olga RussakovskyCVPR 2021
Related papers
- Benign Shortcut for Debiasing: Fair Visual Recognition via Intervention with Shortcut FeaturesYi Zhang, Jitao Sang, Junyang Wang, Dongmei Jiang et al.ACM MM 2023 · 9 citations
- Sustaining Fairness via Incremental LearningSomnath Basu Roy Chowdhury, Snigdha ChaturvediAAAI 2023 · 6 citations
- Fairness Shields: Safeguarding against Biased Decision MakersFilip Cano, Thomas A. Henzinger, Bettina Könighofer, Konstantin Kueffner et al.AAAI 2025
- Gradient Based Activations for Accurate Bias-Free LearningVinod K. Kurmi, Rishabh Sharma, Yash Vardhan Sharma, Vinay P. NamboodiriAAAI 2022 · 3 citations
- Learning for Counterfactual Fairness from Observational DataJing Ma, Ruocheng Guo, Aidong Zhang, Jundong LiKDD 2023 · 9 citations
