Silencing Empowerment, Allowing Bigotry: Auditing the Moderation of Hate Speech on Twitch
Prarabdh Shukla, Wei Yin Chong, Yash Patel, Brennan Schaffner, Danish Pruthi, Arjun Nitin Bhagoji
摘要
Warning: This paper contains content that may be offensive or upsetting. To meet the demands of content moderation, online platforms have resorted to automated systems. Newer forms of real-time engagement (e.g., users commenting on live streams) on platforms like Twitch exert additional pressures on the latency expected of such moderation systems. Despite their prevalence, relatively little is known about the effectiveness of these systems. In this paper, we conduct an audit of Twitch's automated moderation tool (AutoMod) to investigate its effectiveness in flagging hateful content. For our audit, we create streaming accounts to act as siloed test beds, and interface with the live chat using Twitch's APIs to send over 107, 000 comments collated from 4 datasets. We measure AutoMod's accuracy in flagging blatantly hateful content containing misogyny, racism, ableism and homophobia. Our experiments reveal that a large fraction of hateful messages, up to 94% on some datasets, bypass moderation. Contextual addition of slurs to these messages results in 100% removal, revealing AutoMod's reliance on slurs as a moderation signal. We also find that contrary to Twitch's community guidelines, AutoMod blocks up to 89.5% of benign examples that use sensitive words in pedagogical or empowering contexts. Overall, our audit points to large gaps in AutoMod's capabilities and underscores the importance for such systems to understand context effectively. 1 * Equal contribution. 1 Our code and data can be found at https://github. com/weiyinc11/HateSpeechModerationTwitch
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning AttacksAmbra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski 等USENIX Security 2019 · 被引用 466 次
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
- Automated Content Moderation Increases Adherence to Community GuidelinesManoel Horta Ribeiro, Justin Cheng, Robert WestWWW 2023 · 被引用 53 次
- An Image of Society: Gender and Racial Representation and Impact in Image Search Results for OccupationsDanaë Metaxa, Michelle A. Gan, Su Goh, Jeff T. Hancock 等CSCW 2021 · 被引用 52 次
- Hate Raids on Twitch: Echoes of the Past, New Modalities, and Implications for Platform GovernanceCatherine Han, Joseph Seering, Deepak Kumar, Jeffrey T. Hancock 等CSCW 2023 · 被引用 43 次
相关 Paper
- Analyzing Norm Violations in Live-Stream ChatJihyung Moon, Dong-Ho Lee, Hyundong Cho, Woojeong Jin 等EMNLP 2023 · 被引用 5 次
- Hate Raids on Twitch: Understanding Real-Time Human-Bot Coordinated Attacks in Live Streaming CommunitiesJie Cai, Sagnik Chowdhury, Hongyang Zhou, Donghee Yvette WohnCSCW 2023 · 被引用 17 次
- Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic VariationsDavid Hartmann, Amin Oueslati, Dimitri Staufer, Lena Pohlmann 等CHI 2025 · 被引用 37 次
- Coordination and Collaboration: How do Volunteer Moderators Work as a Team in Live Streaming Communities?Jie Cai, Donghee Yvette WohnCHI 2022 · 被引用 34 次
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek 等EMNLP 2024 · 被引用 4 次
