Bias Mitigation for Toxicity Detection via Sequential Decisions
Lu Cheng, Ahmadreza Mosallanezhad, Yasin N. Silva, Deborah L. Hall, Huan Liu
Abstract
Increased social media use has contributed to the greater prevalence of abusive, rude, and offensive textual comments. Machine learning models have been developed to detect toxic comments online, yet these models tend to show biases against users with marginalized or minority identities (e.g., females and African Americans). Established research in debiasing toxicity classifiers often (1) takes a static or batch approach, assuming that all information is available and then making a one-time decision; and (2) uses a generic strategy to mitigate different biases (e.g., gender and racial biases) that assumes the biases are independent of one another. However, in real scenarios, the input typically arrives as a sequence of comments/words over time instead of all at once. Thus, decisions based on partial information must be made while additional input is arriving. Moreover, social bias is complex by nature. Each type of bias is defined within its unique context, which, consistent with intersectionality theory within the social sciences, might be correlated with the contexts of other forms of bias. In this work, we consider debiasing toxicity detection as a sequential decision-making process where different biases can be interdependent. In particular, we study debiasing toxicity detection with two aims: (1) to examine whether different biases tend to correlate with each other; and (2) to investigate how to jointly mitigate these correlated biases in an interactive manner to minimize the total amount of bias. At the core of our approach is a framework built upon theories of sequential Markov Decision Processes that seeks to maximize the prediction accuracy and minimize the bias measures tailored to individual biases. Evaluations on two benchmark datasets empirically validate the hypothesis that biases tend to be correlated and corroborate the effectiveness of the proposed sequential debiasing strategy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6de88d5-3c7b-4e40-82e6-fa16f0b5c537Cited by top-tier papers2
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius et al.KDD 2024 · 7 citations
- Data Caricatures: On the Representation of African American Language in Pretraining CorporaNicholas Deas, Blake Vente, Amith Ananthram, Jessica Grieser et al.ACL 2025
Builds on4
- Demographics Should Not Be the Reason of Toxicity: Mitigating Discrimination in Text Classifications with Instance WeightingGuanhua Zhang, Bing Bai, Junqi Zhang, Kun Bai et al.ACL 2020 · 56 citations
- Toxicity Detection: Does Context Really Matter?John Pavlopoulos, Jeffrey Sorensen, Lucas Dixon, Nithum Thain et al.ACL 2020 · 11 citations
- Mitigating Bias in Session-based Cyberbullying Detection: A Non-Compromising ApproachLu Cheng, Ahmadreza Mosallanezhad, Yasin N. Silva, Deborah L. Hall et al.ACL 2021
- HateCheck: Functional Tests for Hate Speech Detection ModelsPaul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem et al.ACL 2021
Related papers
- Evaluating Short-Term Temporal Fluctuations of Social Biases in Social Media Data and Masked Language ModelsYi Zhou, Danushka Bollegala, José Camacho-ColladosEMNLP 2024 · 3 citations
- MABR: Multilayer Adversarial Bias Removal Without Prior Bias KnowledgeMaxwell J. Yin, Boyu Wang, Charles LingAAAI 2025 · 1 citation
- Centering the Margins: Outlier-Based Identification of Harmed Populations in Toxicity DetectionVyoma Raman, Eve Fleisig, Dan KleinEMNLP 2023 · 1 citation
- A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector SpacesAnne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan VulicAAAI 2020 · 68 citations
- Same Same, But Different: Conditional Multi-Task Learning for Demographic-Specific Toxicity DetectionSoumyajit Gupta, Sooyong Lee, Maria De-Arteaga, Matthew LeaseWWW 2023 · 17 citations
