Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity Annotation
Nitesh Goyal, Ian D. Kivlichan, Rachel Rosen, Lucy Vasserman
Abstract
Machine learning models are commonly used to detect toxicity in online conversations. These models are trained on datasets annotated by human raters. We explore how raters' self-described identities impact how they annotate toxicity in online comments. We first define the concept of specialized rater pools: rater pools formed based on raters' self-described identities, rather than at random. We formed three such rater pools for this study-specialized rater pools of raters from the U.S. who identify as African American, LGBTQ, and those who identify as neither. Each of these rater pools annotated the same set of comments, which contains many references to these identity groups. We found that rater identity is a statistically significant factor in how raters will annotate toxicity for identity-related annotations. Using preliminary content analysis, we examined the comments with the most disagreement between rater pools and found nuanced differences in the toxicity annotations. Next, we trained models on the annotations from each of the different rater pools, and compared the scores of these models on comments from several test sets. Finally, we discuss how using raters that self-identify with the subjects of comments can create more inclusive machine learning models, and provide more nuanced ratings than those by random raters. Please be advised that this work contains examples of toxic and offensive content.
• Human-centered computing → Empirical studies in HCI.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4cddf78f-6898-4074-8f5f-547d4195d686Cited by top-tier papers27
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Synthetic Data Generation with Large Language Models for Text Classification: Potential and LimitationsZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ming YinEMNLP 2023 · 102 citations
- "It is currently hodgepodge": Examining AI/ML Practitioners' Challenges during Co-production of Responsible AI ValuesRama Adithya Varanasi, Nitesh GoyalCHI 2023 · 52 citations
- A hunt for the Snark: Annotator Diversity in Data PracticesShivani Kapania, Alex S. Taylor, Ding WangCHI 2023 · 49 citations
- A Scoping Study of Evaluation Practices for Responsible AI Tools: Steps Towards Effectiveness EvaluationsGlen Berman, Nitesh Goyal, Michael MadaioCHI 2024 · 40 citations
Builds on3
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 236 citations
- Annotating Online MisogynyPhiline Zeinert, Nanna Inie, Leon DerczynskiACL 2021
- HateCheck: Functional Tests for Hate Speech Detection ModelsPaul Röttger, Bertie Vidgen, Dong Nguyen, Zeerak Waseem et al.ACL 2021
Related papers
- ModelCitizens: Representing Community Voices in Online SafetyAshima Suvarna, Christina Chance, Karolina Naranjo, Hamid Palangi et al.EMNLP 2025
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 11 citations
- Learning Subjective Label Distributions via Sociocultural DescriptorsMohammed Fayiz Parappan, Ricardo HenaoEMNLP 2025 · 5 citations
- Bias Mitigation for Toxicity Detection via Sequential DecisionsLu Cheng, Ahmadreza Mosallanezhad, Yasin N. Silva, Deborah L. Hall et al.SIGIR 2022 · 8 citations
- Model Positionality and Computational Reflexivity: Promoting Reflexivity in Data ScienceScott Allen Cambo, Darren GergleCHI 2022 · 49 citations
