A Fine-Grained Taxonomy of Replies to Hate Speech
Xinchen Yu, Ashley Zhao, Eduardo Blanco, Lingzi Hong
Abstract
Countering rather than censoring hate speech has emerged as a promising strategy to address hatred. There are many types of counterspeech in user-generated content: addressing the hateful content or its author, generic requests, well-reasoned counter arguments, insults, etc. The effectiveness of counterspeech, which we define as subsequent incivility, depends on these types. In this paper, we present a theoretically grounded taxonomy of replies to hate speech and a new corpus. We work with real, user-generated hate speech and all the replies it elicits rather than replies generated by a third party. Our analyses provide insights into the content real users reply with as well as which replies are empirically most effective. We also experiment with models to characterize the replies to hate speech, thereby opening the door to estimating whether a reply to hate speech will result in further incivility. Error Type % Example Ground Truth Predicted Rhetorical 26 Hate: F**k worthless inbreds who've contributed nothing to society. question Reply: Where are your contributions? I doubt there's any. Author Content Irony 21 Hate: Retarded republicans fear everything. Reply: It's amazing how broken you have to be to believe in their positions as a whole.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9aebd34b-b55e-4e44-9a05-a95d3be0001fCited by top-tier papers1
Ask how each one uses itBuilds on5
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi et al.EMNLP 2021 · 159 citations
- Perverse Downstream Consequences of Debunking: Being Corrected by Another User for Posting False Political News Increases Subsequent Sharing of Low Quality, Partisan, and Toxic Content in a Twitter Field ExperimentMohsen Mosleh, Cameron Martel, Dean Eckles, David G. RandCHI 2021 · 109 citations
- Detecting Attackable Sentences in ArgumentsYohan Jo, Seojin Bang, Emaad A. Manzoor, Eduard H. Hovy et al.EMNLP 2020 · 25 citations
- Pinpointing Fine-Grained Relationships between Hateful Tweets and RepliesAbdullah Albanyan, Eduardo BlancoAAAI 2022 · 9 citations
- Human-in-the-Loop for Data Collection: a Multi-Target Counter Narrative Dataset to Fight Online Hate SpeechMargherita Fanton, Helena Bonaldi, Serra Sinem Tekiroglu, Marco GueriniACL 2021
Related papers
- Outcome-Constrained Large Language Models for Countering Hate SpeechLingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying SongEMNLP 2024 · 5 citations
- Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech CounteringHelena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio et al.EMNLP 2024 · 2 citations
- Finding Authentic Counterhate Arguments: A Case Study with Public FiguresAbdullah Albanyan, Ahmed Hassan, Eduardo BlancoEMNLP 2023 · 1 citation
- Counterspeakers' Perspectives: Unveiling Barriers and AI Needs in the Fight against Online HateJimin Mun, Cathy Buerger, Jenny T. Liang, Joshua Garland et al.CHI 2024 · 12 citations
- Generating Counter Narratives against Online Hate Speech: Data and StrategiesSerra Sinem Tekiroglu, Yi-Ling Chung, Marco GueriniACL 2020 · 13 citations
