Generating Counter Narratives against Online Hate Speech: Data and Strategies
Serra Sinem Tekiroglu, Yi-Ling Chung, Marco Guerini
Abstract
Recently research has started focusing on avoiding undesired effects that come with content moderation, such as censorship and overblocking, when dealing with hatred online. The core idea is to directly intervene in the discussion with textual responses that are meant to counter the hate content and prevent it from further spreading. Accordingly, automation strategies, such as natural language generation, are beginning to be investigated. Still, they suffer from the lack of sufficient amount of quality data and tend to produce generic/repetitive responses. Being aware of the aforementioned limitations, we present a study on how to collect responses to hate effectively, employing large scale unsupervised language models such as GPT-2 for the generation of silver data, and the best annotation strategies/neural architectures that can be used for data filtering before expert validation/postediting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57ed5dd9-95bc-401d-999f-9f21fd2eec50Cited by top-tier papers16
- "Short is the Road that Leads from Fear to Hate": Fear Speech in Indian WhatsApp GroupsPunyajoy Saha, Binny Mathew, Kiran Garimella, Animesh MukherjeeWWW 2021 · 66 citations
- Reinforcement Learning-based Counter-Misinformation Response Generation: A Case Study of COVID-19 Vaccine MisinformationBing He, Mustaque Ahamad, Srijan KumarWWW 2023 · 62 citations
- CoAnnotating: Uncertainty-Guided Work Allocation between Human and Large Language Models for Data AnnotationMinzhi Li, Taiwei Shi, Caleb Ziems, Min-Yen Kan et al.EMNLP 2023 · 33 citations
- Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and BenchmarksJunyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min et al.ACL 2023 · 25 citations
- Human-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech CounteringHelena Bonaldi, Sara Dellantonio, Serra Sinem Tekiroglu, Marco GueriniEMNLP 2022 · 18 citations
Builds on3
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
Related papers
- Outcome-Constrained Large Language Models for Countering Hate SpeechLingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying SongEMNLP 2024 · 5 citations
- Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech CounteringHelena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio et al.EMNLP 2024 · 2 citations
- Human-in-the-Loop for Data Collection: a Multi-Target Counter Narrative Dataset to Fight Online Hate SpeechMargherita Fanton, Helena Bonaldi, Serra Sinem Tekiroglu, Marco GueriniACL 2021
- A Fine-Grained Taxonomy of Replies to Hate SpeechXinchen Yu, Ashley Zhao, Eduardo Blanco, Lingzi HongEMNLP 2023 · 1 citation
- Evaluation and Facilitation of Online Discussions in the LLM Era: A SurveyKaterina Korre, Dimitris Tsirmpas, Nikos Gkoumas, Emma Cabalé et al.EMNLP 2025
