Human-in-the-Loop for Data Collection: a Multi-Target Counter Narrative Dataset to Fight Online Hate Speech
Margherita Fanton, Helena Bonaldi, Serra Sinem Tekiroglu, Marco Guerini
Abstract
Undermining the impact of hateful content with informed and non-aggressive responses, called counter narratives, has emerged as a possible solution for having healthier online communities. Thus, some NLP studies have started addressing the task of counter narrative generation. Although such studies have made an effort to build hate speech / counter narrative (HS/CN) datasets for neural generation, they fall short in reaching either highquality and/or high-quantity. In this paper, we propose a novel human-in-the-loop data collection methodology in which a generative language model is refined iteratively by using its own data from the previous loops to generate new training samples that experts review and/or post-edit. Our experiments comprised several loops including dynamic variations. Results show that the methodology is scalable and facilitates diverse, novel, and cost-effective data collection. To our knowledge, the resulting dataset is the only expertbased multi-target HS/CN dataset available to the community.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow InstructionsFederico Bianchi, Mirac Suzgun, Giuseppe Attanasio, Paul Röttger et al.ICLR 2024 · 373 citations
- Human-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech CounteringHelena Bonaldi, Sara Dellantonio, Serra Sinem Tekiroglu, Marco GueriniEMNLP 2022 · 18 citations
- Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech GenerationRishabh Gupta, Shaily Desai, Manvi Goel, Anil Bandhakavi et al.ACL 2023 · 10 citations
- LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured DataChuxuan Hu, Austin Peters, Daniel KangVLDB 2025 · 7 citations
- Targeted Data Generation: Finding and Fixing Model WeaknessesZexue He, Marco Túlio Ribeiro, Fereshte KhaniACL 2023 · 6 citations
Builds on4
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Social Bias Frames: Reasoning about Social and Power Implications of LanguageMaarten Sap, Saadia Gabriel, Lianhui Qin, Dan Jurafsky et al.ACL 2020 · 16 citations
- Generating Counter Narratives against Online Hate Speech: Data and StrategiesSerra Sinem Tekiroglu, Yi-Ling Chung, Marco GueriniACL 2020 · 13 citations
- Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe KielaACL 2021
Related papers
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Fact-based Counter Narrative Generation to Combat Hate SpeechBrian Wilk, Homaira Huda Shomee, Suman Kalyan Maity, Sourav MedyaWWW 2025 · 6 citations
- Outcome-Constrained Large Language Models for Countering Hate SpeechLingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying SongEMNLP 2024 · 5 citations
- PerspectiveMod: A Perspectivist Resource for Deliberative ModerationEva Maria Vecchi, Neele Falk, Carlotta Quensel, Iman Jundi et al.EMNLP 2025
- Improving Question Answering Model Robustness with Synthetic Adversarial Data GenerationMax Bartolo, Tristan Thrush, Robin Jia, Sebastian Riedel et al.EMNLP 2021 · 68 citations
