Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
Tharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur R. KhudaBukhsh
Abstract
This paper discusses and contains content that is offensive or disturbing. Offensive speech detection is a key component of content moderation. However, what is offensive can be highly subjective. This paper investigates how machine and human moderators disagree on what is offensive when it comes to real-world social web political discourse. We show that (1) there is extensive disagreement among the moderators (humans and machines); and (2) human and large-language-model classifiers are unable to predict how other human raters will respond, based on their political leanings. For (1), we conduct a noise audit at an unprecedented scale that combines both machine and human responses. For (2), we introduce a firstof-its-kind dataset 1 of vicarious offense. Our noise audit reveals that moderation outcomes vary wildly across different machine moderators. Our experiments with human moderators suggest that political leanings combined with sensitive issues affect both first-person and vicarious offense.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd9a08c2-ee1f-49b4-8c91-027f3b60807dCited by top-tier papers5
- ARTICLE: Annotator Reliability Through In-Context LearningSujan Dutta, Deepak Pandita, Tharindu Cyril Weerasooriya, Marcos Zampieri et al.AAAI 2025 · 7 citations
- Hate Personified: Investigating the role of LLMs in content moderationSarah Masud, Sahajpreet Singh, Viktor Hangya, Alexander Fraser et al.EMNLP 2024 · 6 citations
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 2 citations
- What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency in Closed Domains Via Adversarial NudgeArka Dutta, Sujan Dutta, Rijul Magu, Soumyajit Datta et al.ACL 2026
- Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope SpeechJonathan Pofcher, Christopher M. Homan, Randall Sell, Ashiqur R. KhudaBukhshEMNLP 2025
Builds on5
- "Short is the Road that Leads from Fear to Hate": Fear Speech in Indian WhatsApp GroupsPunyajoy Saha, Binny Mathew, Kiran Garimella, Animesh MukherjeeWWW 2021 · 66 citations
- Pre-train or Annotate? Domain Adaptation with a Constrained BudgetFan Bai, Alan Ritter, Wei XuEMNLP 2021 · 25 citations
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky et al.ACL 2020 · 16 citations
- Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level LearningTharindu Cyril Weerasooriya, Sarah Luger, Saloni Poddar, Ashiqur R. KhudaBukhsh et al.ACL 2023 · 2 citations
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini et al.EMNLP 2021 · 2 citations
Related papers
- Ruddit: Norms of Offensiveness for English Reddit CommentsRishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis et al.ACL 2021
- He said "who's gonna take care of your children when you are at ACL?": Reported Sexist Acts are Not SexistPatricia Chiril, Véronique Moriceau, Farah Benamara, Alda Mari et al.ACL 2020 · 19 citations
- Please note that I'm just an AI: Analysis of Behavior Patterns of LLMs in (Non-)offensive Speech IdentificationEsra Dönmez, Thang Vu, Agnieszka FalenskaEMNLP 2024 · 1 citation
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale et al.ACL 2025 · 12 citations
- Your spouse needs professional help: Determining the Contextual Appropriateness of Messages through Modeling Social RelationshipsDavid Jurgens, Agrima Seth, Jackson Sargent, Athena Aghighi et al.ACL 2023 · 4 citations
