A Fine-Grained Taxonomy of Replies to Hate Speech
Xinchen Yu, Ashley Zhao, Eduardo Blanco, Lingzi Hong
摘要
Countering rather than censoring hate speech has emerged as a promising strategy to address hatred. There are many types of counterspeech in user-generated content: addressing the hateful content or its author, generic requests, well-reasoned counter arguments, insults, etc. The effectiveness of counterspeech, which we define as subsequent incivility, depends on these types. In this paper, we present a theoretically grounded taxonomy of replies to hate speech and a new corpus. We work with real, user-generated hate speech and all the replies it elicits rather than replies generated by a third party. Our analyses provide insights into the content real users reply with as well as which replies are empirically most effective. We also experiment with models to characterize the replies to hate speech, thereby opening the door to estimating whether a reply to hate speech will result in further incivility. Error Type % Example Ground Truth Predicted Rhetorical 26 Hate: F**k worthless inbreds who've contributed nothing to society. question Reply: Where are your contributions? I doubt there's any. Author Content Irony 21 Hate: Retarded republicans fear everything. Reply: It's amazing how broken you have to be to believe in their positions as a whole.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
- Perverse Downstream Consequences of Debunking: Being Corrected by Another User for Posting False Political News Increases Subsequent Sharing of Low Quality, Partisan, and Toxic Content in a Twitter Field ExperimentMohsen Mosleh, Cameron Martel, Dean Eckles, David G. RandCHI 2021 · 被引用 109 次
- Detecting Attackable Sentences in ArgumentsYohan Jo, Seojin Bang, Emaad A. Manzoor, Eduard H. Hovy 等EMNLP 2020 · 被引用 25 次
- Pinpointing Fine-Grained Relationships between Hateful Tweets and RepliesAbdullah Albanyan, Eduardo BlancoAAAI 2022 · 被引用 9 次
- Human-in-the-Loop for Data Collection: a Multi-Target Counter Narrative Dataset to Fight Online Hate SpeechMargherita Fanton, Helena Bonaldi, Serra Sinem Tekiroglu, Marco GueriniACL 2021
相关 Paper
- Outcome-Constrained Large Language Models for Countering Hate SpeechLingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying SongEMNLP 2024 · 被引用 5 次
- Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech CounteringHelena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio 等EMNLP 2024 · 被引用 2 次
- Finding Authentic Counterhate Arguments: A Case Study with Public FiguresAbdullah Albanyan, Ahmed Hassan, Eduardo BlancoEMNLP 2023 · 被引用 1 次
- Counterspeakers' Perspectives: Unveiling Barriers and AI Needs in the Fight against Online HateJimin Mun, Cathy Buerger, Jenny T. Liang, Joshua Garland 等CHI 2024 · 被引用 12 次
- Generating Counter Narratives against Online Hate Speech: Data and StrategiesSerra Sinem Tekiroglu, Yi-Ling Chung, Marco GueriniACL 2020 · 被引用 13 次
