BiasWipe: Mitigating Unintended Bias in Text Classifiers through Model Interpretability
Mamta Mamta, Rishikant Chigrupaatii, Asif Ekbal
摘要
Toxic content detection plays a vital role in addressing the misuse of social media platforms to harm people or groups due to their race, gender or ethnicity. However, due to the nature of the datasets, systems develop an unintended bias due to the over-generalization of the model to the training data. This compromises the fairness of the systems, which can impact certain groups due to their race, gender, etc. Existing methods mitigate bias using data augmentation, adversarial learning, etc., which require re-training and adding extra parameters to the model. In this work, we present a robust and generalizable technique BiasWipe to mitigate unintended bias in language models. BiasWipe utilizes model interpretability using Shapley values, which achieve fairness by pruning the neuron weights responsible for unintended bias. It first identifies the neuron weights responsible for unintended bias and then achieves fairness by pruning them without loss of original performance. It does not require re-training or adding extra parameters to the model. To show the effectiveness of our proposed technique for bias unlearning, we perform extensive experiments for Toxic content detection for BERT, RoBERTa, and GPT models. 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?Yifan Wang, Mayank Jobanputra, Ji-Ung Lee, Soyoung Oh 等ICLR 2026 · 被引用 3 次
- Debiasing the Fine-Grained Classification Task in LLMs with Bias-Aware PEFTDaiying Zhao, Xinyu Yang, Hang ChenACL 2025 · 被引用 1 次
它引用的顶会 Paper9
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 被引用 195 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- Transformer Feed-Forward Layers Are Key-Value MemoriesMor Geva, Roei Schuster, Jonathan Berant, Omer LevyEMNLP 2021 · 被引用 33 次
- Finding Skill Neurons in Pre-trained Transformer-based Language ModelsXiaozhi Wang, Kaiyue Wen, Zhengyan Zhang, Lei Hou 等EMNLP 2022 · 被引用 19 次
- Elevating Code-mixed Text Handling through Auditory Information of WordsMamta, Zishan Ahmad, Asif EkbalEMNLP 2023 · 被引用 6 次
相关 Paper
- The Devil is in the Neurons: Interpreting and Mitigating Social Biases in Language ModelsYan Liu, Yu Liu, Xiaokang Chen, Pin-Yu Chen 等ICLR 2024 · 被引用 32 次
- Mitigating Biases for Instruction-following Language Models via Bias Neurons EliminationNakyeong Yang, Taegwan Kang, Stanley Jungkyu Choi, Honglak Lee 等ACL 2024
- Leashing the Inner Demons: Self-Detoxification for Language ModelsCanwen Xu, Zexue He, Zhankui He, Julian J. McAuleyAAAI 2022 · 被引用 30 次
- ConceptPrune: Concept Editing in Diffusion Models via Skilled Neuron PruningRuchika Chavhan, Da Li, Timothy M. HospedalesICLR 2025 · 被引用 2 次
- Self-Detoxifying Language Models via Toxification ReversalChak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang 等EMNLP 2023 · 被引用 12 次
