Bias Mitigation for Toxicity Detection via Sequential Decisions
Lu Cheng, Ahmadreza Mosallanezhad, Yasin N. Silva, Deborah L. Hall, Huan Liu
摘要
Increased social media use has contributed to the greater prevalence of abusive, rude, and offensive textual comments. Machine learning models have been developed to detect toxic comments online, yet these models tend to show biases against users with marginalized or minority identities (e.g., females and African Americans). Established research in debiasing toxicity classifiers often (1) takes a static or batch approach, assuming that all information is available and then making a one-time decision; and (2) uses a generic strategy to mitigate different biases (e.g., gender and racial biases) that assumes the biases are independent of one another. However, in real scenarios, the input typically arrives as a sequence of comments/words over time instead of all at once. Thus, decisions based on partial information must be made while additional input is arriving. Moreover, social bias is complex by nature. Each type of bias is defined within its unique context, which, consistent with intersectionality theory within the social sciences, might be correlated with the contexts of other forms of bias. In this work, we consider debiasing toxicity detection as a sequential decision-making process where different biases can be interdependent. In particular, we study debiasing toxicity detection with two aims: (1) to examine whether different biases tend to correlate with each other; and (2) to investigate how to jointly mitigate these correlated biases in an interactive manner to minimize the total amount of bias. At the core of our approach is a framework built upon theories of sequential Markov Decision Processes that seeks to maximize the prediction accuracy and minimize the bias measures tailored to individual biases. Evaluations on two benchmark datasets empirically validate the hypothesis that biases tend to be correlated and corroborate the effectiveness of the proposed sequential debiasing strategy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius 等KDD 2024 · 被引用 7 次
- Data Caricatures: On the Representation of African American Language in Pretraining CorporaNicholas Deas, Blake Vente, Amith Ananthram, Jessica Grieser 等ACL 2025
它引用的顶会 Paper4
- Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance WeightingGuanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai 等ACL 2020 · 被引用 56 次
- Toxicity Detection: Does Context Really Matter?John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain 等ACL 2020 · 被引用 11 次
- Mitigating Bias in Session-based Cyberbullying Detection: A Non-Compromising ApproachLu Cheng, Ahmadreza Mosallanezhad, Yasin N. Silva, Deborah L. Hall 等ACL 2021
- HateCheck: Functional Tests for Hate Speech Detection ModelsPaul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem 等ACL 2021
相关 Paper
- Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language ModelsYi Zhou, Danushka Bollegala, José Camacho-ColladosEMNLP 2024 · 被引用 3 次
- MABR: Multilayer Adversarial Bias Removal Without Prior Bias KnowledgeMaxwell J. Yin, Boyu Wang, Charles LingAAAI 2025 · 被引用 1 次
- Centering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity DetectionVyoma Raman, Eve Fleisig, Dan KleinEMNLP 2023 · 被引用 1 次
- A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector SpacesAnne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan VulicAAAI 2020 · 被引用 68 次
- Same Same, But Different: Conditional Multi-Task Learning for Demographic-Specific Toxicity DetectionSoumyajit Gupta, Sooyong Lee, Maria De-Arteaga, Matthew LeaseWWW 2023 · 被引用 17 次
