Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
David Hartmann, Amin Oueslati, Dimitri Staufer, Lena Pohlmann, Simon Munzert, Hendrik Heuer
Abstract
Commercial content moderation APIs are marketed as scalable solutions to combat online hate speech. However, the reliance on these APIs risks both silencing legitimate speech, called over-moderation, and failing to protect online platforms from harmful speech, known as under-moderation. To assess such risks, this paper introduces a framework for auditing black-box NLP systems. Using the framework, we systematically evaluate five widely used commercial content moderation APIs. Analyzing five million queries based on four datasets, we find that APIs frequently rely on group identity terms, such as "black", to predict hate speech. While OpenAI's and Amazon's services perform slightly better, all providers undermoderate implicit hate speech, which uses codified messages, especially against LGBTQIA+ individuals. Simultaneously, they overmoderate counter-speech, reclaimed slurs and content related to Black, LGBTQIA+, Jewish, and Muslim people. We recommend that API providers offer better guidance on API implementation and threshold setting and more transparency on their APIs' limitations. Warning: This paper contains offensive and hateful terms and concepts. We have chosen to reproduce these terms for reasons of transparency.
• Human-centered computing → Empirical studies in HCI; Empirical studies in collaborative and social computing; • General and reference → Measurement; • Social and professional topics → Hate speech.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 008d9e79-e21a-4299-972e-f1d21bd2a71bCited by top-tier papers6
- HateDay: Insights from a Global Hate Speech Dataset Representative of a Day on TwitterManuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A. Hale et al.ACL 2025 · 12 citations
- HateBuffer: Safeguarding Content Moderators' Mental Well-Being through Hate Speech Content ModificationSubin Park, Jeonghyun Kim, Jeanne Choi, Joseph Seering et al.CSCW 2025 · 4 citations
- Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on TwitchPrarabdh Shukla, Wei Yin Chong, Yash Patel, Brennan Schaffner et al.ACL 2025 · 3 citations
- Echoes of Norms: Investigating Counterspeech Bots' Influence on Bystanders in Online CommunitiesMengyao Wang, Shuai Ma, Nuo Li, Peng Zhang et al.CHI 2026 · 1 citation
- Take the Power Back: Screen-Based Personal Moderation Against Hate Speech on InstagramAnna Ricarda Luther, Hendrik Heuer, Sebastian Haunss, Stephanie Geise et al.CHI 2026 · 1 citation
Builds on15
- Problematic Machine Behavior: A Systematic Literature Review of Algorithm AuditsJack BandyCSCW 2021 · 190 citations
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi et al.EMNLP 2021 · 159 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- "It's common and a part of being a content creator": Understanding How Creators Experience and Cope with Hate and Harassment OnlineKurt Thomas, Patrick Gage Kelley, Sunny Consolvo, Patrawat Samermit et al.CHI 2022 · 63 citations
- Decolonizing Content Moderation: Does Uniform Global Community Standard Resemble Utopian Equality or Western Power Hegemony?Farhana Shahid, Aditya VashisthaCHI 2023 · 63 citations
Related papers
- There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language ModelsFriedemann Lipphardt, Moonis Ali, Martin Banzer, Anja Feldmann et al.NDSS 2026 · 1 citation
- Please note that I'm just an AI: Analysis of Behavior Patterns of LLMs in (Non-)offensive Speech IdentificationEsra Dönmez, Thang Vu, Agnieszka FalenskaEMNLP 2024 · 1 citation
- "Ignorance is not Bliss": Designing Personalized Moderation to Address Ableist Hate on Social MediaSharon Heung, Lucy Jiang, Shiri Azenkot, Aditya VashisthaCHI 2025 · 14 citations
- NaijaHate: Evaluating Hate Speech Detection on Nigerian Twitter Using Representative DataManuel Tonneau, Pedro Vitor Quinta de Castro, Karim Lasri, Ibrahim Farouq et al.ACL 2024
- Vicarious Offense and Noise Audit of Offensive Speech Classifiers: Unifying Human and Machine Disagreement on What is OffensiveTharindu Cyril Weerasooriya, Sujan Dutta, Tharindu Ranasinghe, Marcos Zampieri et al.EMNLP 2023 · 12 citations
