Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance Weighting
Guanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai, Conghui Zhu, Tiejun Zhao
摘要
With the recent proliferation of the use of text classifications, researchers have found that there are certain unintended biases in text classification datasets. For example, texts containing some demographic identity-terms (e.g., "gay", "black") are more likely to be abusive in existing abusive language detection datasets. As a result, models trained with these datasets may consider sentences like "She makes me happy to be gay" as abusive simply because of the word "gay." In this paper, we formalize the unintended biases in text classification datasets as a kind of selection bias from the non-discrimination distribution to the discrimination distribution. Based on this formalization, we further propose a model-agnostic debiasing training framework by recovering the non-discrimination distribution using instance weighting, which does not require any extra resources or annotations apart from a pre-defined set of demographic identity-terms. Experiments demonstrate that our method can effectively alleviate the impacts of the unintended biases without significantly hurting models' generalization ability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Why Attentions May Not Be Interpretable?Bing Bai, Jian Liang, Guanhua Zhang, Hao Li 等KDD 2021 · 被引用 51 次
- Fairness ReprogrammingGuanhua Zhang, Yihua Zhang, Yang Zhang, Wenqi Fan 等NeurIPS 2022 · 被引用 46 次
- A Training-Free Debiasing Framework with Counterfactual Reasoning for Conversational Emotion DetectionGeng Tu, Ran Jing, Bin Liang, Min Yang 等EMNLP 2023 · 被引用 9 次
- Out-of-Distribution Detection via Conditional Kernel Independence ModelYu Wang, Jingjing Zou, Jingyang Lin, Qing Ling 等NeurIPS 2022 · 被引用 9 次
- Can We Improve Model Robustness through Secondary Attribute Counterfactuals?Ananth Balashankar, Xuezhi Wang, Ben Packer, Nithum Thain 等EMNLP 2021 · 被引用 9 次
相关 Paper
- Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual ClassifiersQuentin Guimard, Moreno D'Incà, Massimiliano Mancini, Elisa RicciCVPR 2025
- Mitigating Bias in Session-based Cyberbullying Detection: A Non-Compromising ApproachLu Cheng, Ahmadreza Mosallanezhad, Yasin N. Silva, Deborah L. Hall 等ACL 2021
- Counterfactual Inference for Text Classification DebiasingChen Qian, Fuli Feng, Lijie Wen, Chunping Ma 等ACL 2021
- A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector SpacesAnne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan VulicAAAI 2020 · 被引用 68 次
- Debiasing Pretrained Text Encoders by Paying Attention to Paying AttentionYacine Gaci, Boualem Benatallah, Fabio Casati, Khalid BenabdeslemEMNLP 2022 · 被引用 12 次
