Outcome-Constrained Large Language Models for Countering Hate Speech
Lingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying Song
Abstract
Automatic counterspeech generation methods have been developed to assist efforts in combating hate speech. Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven. However, the real impact of counterspeech in online environments is seldom considered. This study aims to develop methods for generating counterspeech constrained by conversation outcomes and evaluate their effectiveness. We experiment with large language models (LLMs) to incorporate into the text generation process two desired conversation outcomes: low conversation incivility and nonhateful hater reentry. Specifically, we experiment with instruction prompts, LLM finetuning, and LLM reinforcement learning (RL). Evaluation results show that our methods effectively steer the generation of counterspeech towards the desired outcomes. Our analyses, however, show that there are differences in the quality and style depending on the model.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f7cf580-1c13-46cd-bf0c-efd7d4ffec84Cited by top-tier papers2
- Echoes of Norms: Investigating Counterspeech Bots' Influence on Bystanders in Online CommunitiesMengyao Wang, Shuai Ma, Nuo Li, Peng Zhang et al.CHI 2026 · 1 citation
- Exploring Selective Avoidance for Online User Behavior Analysis: A Forest of Thought ExplanationXiaohua Wu, Lin Li, Kaize Shi, Xiaohui Tao et al.AAAI 2026
Builds on7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu et al.ACL 2020 · 229 citations
- Controlled Text Generation as Continuous Optimization with Multiple ConstraintsSachin Kumar, Eric Malmi, Aliaksei Severyn, Yulia TsvetkovNeurIPS 2021 · 91 citations
- Generating Counter Narratives against Online Hate Speech: Data and StrategiesSerra Sinem Tekiroglu, Yi-Ling Chung, Marco GueriniACL 2020 · 13 citations
- Counterspeakers' Perspectives: Unveiling Barriers and AI Needs in the Fight against Online HateJimin Mun, Cathy Buerger, Jenny T. Liang, Joshua Garland et al.CHI 2024 · 12 citations
Related papers
- Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech CounteringHelena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio et al.EMNLP 2024 · 2 citations
- LLM-based Rewriting of Inappropriate Argumentation using Reinforcement Learning from Machine FeedbackTimon Ziegenbein, Gabriella Skitalinskaya, Alireza Bayat Makou, Henning WachsmuthACL 2024
- A Fine-Grained Taxonomy of Replies to Hate SpeechXinchen Yu, Ashley Zhao, Eduardo Blanco, Lingzi HongEMNLP 2023 · 1 citation
- F²RL: Factuality and Faithfulness Reinforcement Learning Framework for Claim-Guided Evidence-Supported Counterspeech GenerationHaiyang Wang, Yuchen Pan, Xin Song, Xuechen Zhao et al.EMNLP 2024 · 1 citation
- Comparing human and LLM politeness strategies in free productionHaoran Zhao, Robert D. HawkinsEMNLP 2025 · 2 citations
