Fairness via Representation Neutralization
Mengnan Du, Subhabrata Mukherjee, Guanchu Wang, Ruixiang Tang, Ahmed Hassan Awadallah, Xia Ben Hu
Abstract
Existing bias mitigation methods for DNN models primarily work on learning debiased encoders. This process not only requires a lot of instance-level annotations for sensitive attributes, it also does not guarantee that all fairness sensitive information has been removed from the encoder. To address these limitations, we explore the following research question: Can we reduce the discrimination of DNN models by only debiasing the classification head, even with biased representations as inputs? To this end, we propose a new mitigation technique, namely, Representation Neutralization for Fairness (RNF) that achieves fairness by debiasing only the task-specific classification head of DNN models. To this end, we leverage samples with the same ground-truth label but different sensitive attributes, and use their neutralized representations to train the classification head of the DNN model. The key idea of RNF is to discourage the classification head from capturing undesirable correlation between fairness sensitive information in encoder representations with specific class labels. To address low-resource settings with no access to sensitive attribute annotations, we leverage a bias-amplified model to generate proxy annotations for sensitive attributes. Experimental results over several benchmark datasets demonstrate our RNF framework to effectively reduce discrimination of DNN models with minimal degradation in task-specific performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f363e70-417e-4992-82b8-00bbc11c3e98Cited by top-tier papers10
- Generalized Demographic Parity for Group FairnessZhimeng Jiang, Xiaotian Han, Chao Fan, Fan Yang et al.ICLR 2022 · 71 citations
- Fair Graph DistillationQizhang Feng, Zhimeng Stephen Jiang, Ruiquan Li, Yicheng Wang et al.NeurIPS 2023 · 21 citations
- FAIRER: Fairness as Decision Rationale AlignmentTianlin Li, Qing Guo, Aishan Liu, Mengnan Du et al.ICML 2023 · 20 citations
- Input-agnostic Certified Group Fairness via Gaussian Parameter SmoothingJiayin Jin, Zeru Zhang, Yang Zhou, Lingfei WuICML 2022 · 18 citations
- Combating Unknown Bias with Effective Bias-Conflicting Scoring and Gradient AlignmentBowen Zhao, Chen Chen, Qian-Wei Wang, Anfeng He et al.AAAI 2023 · 11 citations
Builds on7
- Decoupling Representation and Classifier for Long-Tailed RecognitionBingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan et al.ICLR 2020 · 1,496 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Gradient Starvation: A Learning Proclivity in Neural NetworksMohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron C. Courville et al.NeurIPS 2021 · 378 citations
- Invariant Risk Minimization GamesKartik Ahuja, Karthikeyan Shanmugam, Kush R. Varshney, Amit DhurandharICML 2020 · 289 citations
- Interpretations are Useful: Penalizing Explanations to Align Neural Networks with Prior KnowledgeLaura Rieger, Chandan Singh, W. James Murdoch, Bin YuICML 2020 · 249 citations
Related papers
- Towards Debiasing DNN Models from Spurious Feature InfluenceMengnan Du, Ruixiang Tang, Weijie Fu, Xia HuAAAI 2022 · 9 citations
- Mining bias-target Alignment from Voronoi CellsRémi Nahon, Van-Tam Nguyen, Enzo TartaglioneICCV 2023 · 8 citations
- Gradient Based Activations for Accurate Bias-Free LearningVinod K. Kurmi, Rishabh Sharma, Yash Vardhan Sharma, Vinay P. NamboodiriAAAI 2022 · 3 citations
- FairSIN: Achieving Fairness in Graph Neural Networks through Sensitive Information NeutralizationCheng Yang, Jixi Liu, Yunhe Yan, Chuan ShiAAAI 2024 · 38 citations
- Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of SamplesWeiwei Li, Junzhuo Liu, Yuanyuan Ren, Yuchen Zheng et al.CVPR 2025
