MABR: Multilayer Adversarial Bias Removal Without Prior Bias Knowledge
Maxwell J. Yin, Boyu Wang, Charles Ling
摘要
Models trained on real-world data often mirror and exacerbate existing social biases. Traditional methods for mitigating these biases typically require prior knowledge of the specific biases to be addressed, and the social groups associated with each instance. In this paper, we introduce a novel adversarial training strategy that operates withour relying on prior bias-type knowledge (e.g., gender or racial bias) and protected attribute labels. Our approach dynamically identifies biases during model training by utilizing auxiliary bias detector. These detected biases are simultaneously mitigated through adversarial training. Crucially, we implement these bias detectors at various levels of the feature maps of the main model, enabling the detection of a broader and more nuanced range of bias features. Through experiments on racial and gender biases in sentiment and occupation classification tasks, our method effectively reduces social biases without the need for demographic annotations. Moreover, our approach not only matches but often surpasses the efficacy of methods that require detailed demographic insights, marking a significant advancement in bias mitigation techniques.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 被引用 3,729 次
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan 等ICML 2021 · 被引用 683 次
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang 等ICCV 2019 · 被引用 469 次
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee 等NeurIPS 2020 · 被引用 406 次
- Linear Adversarial Concept ErasureShauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan CotterellICML 2022 · 被引用 89 次
相关 Paper
- BLIND: Bias Removal With No DemographicsHadas Orgad, Yonatan BelinkovACL 2023 · 被引用 8 次
- Gradient Based Activations for Accurate Bias-Free LearningVinod K. Kurmi, Rishabh Sharma, Yash Vardhan Sharma, Vinay P. NamboodiriAAAI 2022 · 被引用 3 次
- Fair Attribute Classification Through Latent Space De-BiasingVikram V. Ramaswamy, Sunnie S. Y. Kim, Olga RussakovskyCVPR 2021
- NLPGuard: A Framework for Mitigating the Use of Protected Attributes by NLP ClassifiersSalvatore Greco, Ke Zhou, Licia Capra, Tania Cerquitelli 等CSCW 2024 · 被引用 5 次
- Balancing out Bias: Achieving Fairness Through Balanced TrainingXudong Han, Timothy Baldwin, Trevor CohnEMNLP 2022 · 被引用 17 次
