Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
David Hartmann, Amin Oueslati, Dimitri Staufer, Lena Pohlmann, Simon Munzert, Hendrik Heuer
摘要
Commercial content moderation APIs are marketed as scalable solutions to combat online hate speech. However, the reliance on these APIs risks both silencing legitimate speech, called over-moderation, and failing to protect online platforms from harmful speech, known as under-moderation. To assess such risks, this paper introduces a framework for auditing black-box NLP systems. Using the framework, we systematically evaluate five widely used commercial content moderation APIs. Analyzing five million queries based on four datasets, we find that APIs frequently rely on group identity terms, such as "black", to predict hate speech. While OpenAI's and Amazon's services perform slightly better, all providers undermoderate implicit hate speech, which uses codified messages, especially against LGBTQIA+ individuals. Simultaneously, they overmoderate counter-speech, reclaimed slurs and content related to Black, LGBTQIA+, Jewish, and Muslim people. We recommend that API providers offer better guidance on API implementation and threshold setting and more transparency on their APIs' limitations. Warning: This paper contains offensive and hateful terms and concepts. We have chosen to reproduce these terms for reasons of transparency.
• Human-centered computing → Empirical studies in HCI; Empirical studies in collaborative and social computing; • General and reference → Measurement; • Social and professional topics → Hate speech.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale 等ACL 2025 · 被引用 12 次
- HateBuffer: Safeguarding Content Moderators' Mental Well-Being through Hate Speech Content ModificationSubin Park, Jeonghyun Kim, Jeanne Choi, Joseph Seering 等CSCW 2025 · 被引用 4 次
- Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on TwitchPrarabdh Shukla, Wei Yin Chong, Yash Patel, Brennan Schaffner 等ACL 2025 · 被引用 3 次
- Echoes of Norms: Investigating Counterspeech Bots' Influence on Bystanders in Online CommunitiesMengyao Wang, Shuai Ma, Nuo Li, Peng Zhang 等CHI 2026 · 被引用 1 次
- Take the Power Back: Screen-Based Personal Moderation Against Hate Speech on InstagramAnna Ricarda Luther, Hendrik Heuer, Sebastian Haunss, Stephanie Geise 等CHI 2026 · 被引用 1 次
它引用的顶会 Paper15
- Problematic Machine Behavior: A Systematic Literature Review of Algorithm AuditsJack BandyCSCW 2021 · 被引用 190 次
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- "It's common and a part of being a content creator": Understanding How Creators Experience and Cope with Hate and Harassment OnlineKurt Thomas, Patrick Gage Kelley, Sunny Consolvo, Patrawat Samermit 等CHI 2022 · 被引用 63 次
- Decolonizing Content Moderation: Does Uniform Global Community Standard Resemble Utopian Equality or Western Power Hegemony?Farhana Shahid, Aditya VashisthaCHI 2023 · 被引用 63 次
相关 Paper
- There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language ModelsFriedemann Lipphardt, Moonis Ali, Martin Banzer, Anja Feldmann 等NDSS 2026 · 被引用 1 次
- Please note that I'm just an AI: Analysis of Behavior Patterns of LLMs in (Non-)offensive Speech IdentificationEsra Dönmez, Thang Vu, Agnieszka FalenskaEMNLP 2024 · 被引用 1 次
- "Ignorance is not Bliss": Designing Personalized Moderation to Address Ableist Hate on Social MediaSharon Heung, Lucy Jiang, Shiri Azenkot, Aditya VashisthaCHI 2025 · 被引用 14 次
- NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataManuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq 等ACL 2024
- Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveTharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri 等EMNLP 2023 · 被引用 12 次
