Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is Offensive
Tharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri, Christopher Homan, Ashiqur R. KhudaBukhsh
摘要
This paper discusses and contains content that is offensive or disturbing. Offensive speech detection is a key component of content moderation. However, what is offensive can be highly subjective. This paper investigates how machine and human moderators disagree on what is offensive when it comes to real-world social web political discourse. We show that (1) there is extensive disagreement among the moderators (humans and machines); and (2) human and large-language-model classifiers are unable to predict how other human raters will respond, based on their political leanings. For (1), we conduct a noise audit at an unprecedented scale that combines both machine and human responses. For (2), we introduce a firstof-its-kind dataset 1 of vicarious offense. Our noise audit reveals that moderation outcomes vary wildly across different machine moderators. Our experiments with human moderators suggest that political leanings combined with sensitive issues affect both first-person and vicarious offense.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- ARTICLE: Annotator Reliability Through In-Context LearningSujan Dutta, Deepak Pandita, Tharindu Cyril Weerasooriya, Marcos Zampieri 等AAAI 2025 · 被引用 7 次
- Hate Personified: Investigating the role of LLMs in content moderationSarah Masud, Sahajpreet Singh, Viktor Hangya, Alexander Fraser 等EMNLP 2024 · 被引用 6 次
- Forest vs Tree: The (N, K) Trade-off in Reproducible ML EvaluationDeepak Pandita, Flip Korn, Chris Welty, Christopher M. HomanAAAI 2026 · 被引用 2 次
- What About the Scene With the Hitler Reference? HAUNT: A Framework to Probe LLMs' Self-consistency in Closed Domains Via Adversarial NudgeArka Dutta, Sujan Dutta, Rijul Magu, Soumyajit Datta 等ACL 2026
- Hope vs. Hate: Understanding User Interactions with LGBTQ+ News Content in Mainstream US News Media through the Lens of Hope SpeechJonathan Pofcher, Christopher M. Homan, Randall Sell, Ashiqur R. KhudaBukhshEMNLP 2025
它引用的顶会 Paper5
- "Short is the Road that Leads from Fear to Hate": Fear Speech in Indian WhatsApp GroupsPunyajoy Saha, Binny Mathew, Kiran Garimella, Animesh MukherjeeWWW 2021 · 被引用 66 次
- Pre-train or Annotate? Domain Adaptation with a Constrained BudgetFan Bai, Alan Ritter, Wei XuEMNLP 2021 · 被引用 25 次
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky 等ACL 2020 · 被引用 16 次
- Subjective Crowd Disagreements for Subjective Data: Uncovering Meaningful CrowdOpinion with Population-level LearningTharindu Cyril Weerasooriya, Sarah Luger, Saloni Poddar, Ashiqur R. KhudaBukhsh 等ACL 2023 · 被引用 2 次
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini 等EMNLP 2021 · 被引用 2 次
相关 Paper
- Ruddit: Norms of Offensiveness for English Reddit CommentsRishav Hada, Sohi Sudhir, Pushkar Mishra, Helen Yannakoudakis 等ACL 2021
- He said "who's gonna take care of your children when you are at ACL?": Reported Sexist Acts are Not SexistPatricia Chiril, Véronique Moriceau, Farah Benamara, Alda Mari 等ACL 2020 · 被引用 19 次
- Please note that I'm just an AI: Analysis of Behavior Patterns of LLMs in (Non-)offensive Speech IdentificationEsra Dönmez, Thang Vu, Agnieszka FalenskaEMNLP 2024 · 被引用 1 次
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale 等ACL 2025 · 被引用 12 次
- Your spouse needs professional help: Determining the Contextual Appropriateness of Messages through Modeling Social RelationshipsDavid Jurgens, Agrima Seth, Jackson Sargent, Athena Aghighi 等ACL 2023 · 被引用 4 次
