Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch
Prarabdh Shukla, Wei Yin Chong, Yash Patel, Brennan Schaffner, Danish Pruthi, Arjun Nitin Bhagoji
Abstract
Warning: This paper contains content that may be offensive or upsetting. To meet the demands of content moderation, online platforms have resorted to automated systems. Newer forms of real-time engagement (e.g., users commenting on live streams) on platforms like Twitch exert additional pressures on the latency expected of such moderation systems. Despite their prevalence, relatively little is known about the effectiveness of these systems. In this paper, we conduct an audit of Twitch's automated moderation tool (AutoMod) to investigate its effectiveness in flagging hateful content. For our audit, we create streaming accounts to act as siloed test beds, and interface with the live chat using Twitch's APIs to send over 107, 000 comments collated from 4 datasets. We measure AutoMod's accuracy in flagging blatantly hateful content containing misogyny, racism, ableism and homophobia. Our experiments reveal that a large fraction of hateful messages, up to 94% on some datasets, bypass moderation. Contextual addition of slurs to these messages results in 100% removal, revealing AutoMod's reliance on slurs as a moderation signal. We also find that contrary to Twitch's community guidelines, AutoMod blocks up to 89.5% of benign examples that use sensitive words in pedagogical or empowering contexts. Overall, our audit points to large gaps in AutoMod's capabilities and underscores the importance for such systems to understand context effectively. 1 * Equal contribution. 1 Our code and data can be found at https://github. com/weiyinc11/HateSpeechModerationTwitch
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a90191b5-982a-4bf1-a8e7-43bf716f553cCited by top-tier papers1
Ask how each one uses itBuilds on10
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning AttacksAmbra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski et al.USENIX Security 2019 · 466 citations
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi et al.EMNLP 2021 · 159 citations
- Automated Content Moderation Increases Adherence to Community GuidelinesManoel Horta Ribeiro, Justin Cheng, Robert WestWWW 2023 · 53 citations
- An Image of Society: Gender and Racial Representation and Impact in Image Search Results for OccupationsDanaë Metaxa, Michelle A. Gan, Su Goh, Jeff T. Hancock et al.CSCW 2021 · 52 citations
- Hate Raids on Twitch: Echoes of the Past, New Modalities, and Implications for Platform GovernanceCatherine Han, Joseph Seering, Deepak Kumar, Jeffrey T. Hancock et al.CSCW 2023 · 43 citations
Related papers
- Analyzing Norm Violations in Live-Stream ChatJihyung Moon, Dong-Ho Lee, Hyundong Cho, Woojeong Jin et al.EMNLP 2023 · 5 citations
- Hate Raids on Twitch: Understanding Real-Time Human-Bot Coordinated Attacks in Live Streaming CommunitiesJie Cai, Sagnik Chowdhury, Hongyang Zhou, Donghee Yvette WohnCSCW 2023 · 17 citations
- Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic VariationsDavid Hartmann, Amin Oueslati, Dimitri Staufer, Lena Pohlmann et al.CHI 2025 · 37 citations
- Coordination and Collaboration: How do Volunteer Moderators Work as a Team in Live Streaming Communities?Jie Cai, Donghee Yvette WohnCHI 2022 · 34 citations
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek et al.EMNLP 2024 · 4 citations
