MABR: Multilayer Adversarial Bias Removal Without Prior Bias Knowledge
Maxwell J. Yin, Boyu Wang, Charles Ling
Abstract
Models trained on real-world data often mirror and exacerbate existing social biases. Traditional methods for mitigating these biases typically require prior knowledge of the specific biases to be addressed, and the social groups associated with each instance. In this paper, we introduce a novel adversarial training strategy that operates withour relying on prior bias-type knowledge (e.g., gender or racial bias) and protected attribute labels. Our approach dynamically identifies biases during model training by utilizing auxiliary bias detector. These detected biases are simultaneously mitigated through adversarial training. Crucially, we implement these bias detectors at various levels of the feature maps of the main model, enabling the detection of a broader and more nuanced range of bias features. Through experiments on racial and gender biases in sentiment and occupation classification tasks, our method effectively reduces social biases without the need for demographic annotations. Moreover, our approach not only matches but often surpasses the efficacy of methods that require detailed demographic insights, marking a significant advancement in bias mitigation techniques.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d9d6e2f-59f4-42fe-aac1-875d1cc2c266Builds on13
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee et al.NeurIPS 2020 · 406 citations
- Linear Adversarial Concept ErasureShauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan CotterellICML 2022 · 89 citations
Related papers
- BLIND: Bias Removal With No DemographicsHadas Orgad, Yonatan BelinkovACL 2023 · 8 citations
- Gradient Based Activations for Accurate Bias-Free LearningVinod K. Kurmi, Rishabh Sharma, Yash Vardhan Sharma, Vinay P. NamboodiriAAAI 2022 · 3 citations
- Fair Attribute Classification Through Latent Space De-BiasingVikram V. Ramaswamy, Sunnie S. Y. Kim, Olga RussakovskyCVPR 2021
- NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP ClassifiersSalvatore Greco, Ke Zhou, Licia Capra, Tania Cerquitelli et al.CSCW 2024 · 5 citations
- Balancing out Bias: Achieving Fairness Through Balanced TrainingXudong Han, Timothy Baldwin, Trevor CohnEMNLP 2022 · 17 citations
