Outcome-Constrained Large Language Models for Countering Hate Speech
Lingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying Song
摘要
Automatic counterspeech generation methods have been developed to assist efforts in combating hate speech. Existing research focuses on generating counterspeech with linguistic attributes such as being polite, informative, and intent-driven. However, the real impact of counterspeech in online environments is seldom considered. This study aims to develop methods for generating counterspeech constrained by conversation outcomes and evaluate their effectiveness. We experiment with large language models (LLMs) to incorporate into the text generation process two desired conversation outcomes: low conversation incivility and nonhateful hater reentry. Specifically, we experiment with instruction prompts, LLM finetuning, and LLM reinforcement learning (RL). Evaluation results show that our methods effectively steer the generation of counterspeech towards the desired outcomes. Our analyses, however, show that there are differences in the quality and style depending on the model.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Echoes of Norms: Investigating Counterspeech Bots' Influence on Bystanders in Online CommunitiesMengyao Wang, Shuai Ma, Nuo Li, Peng Zhang 等CHI 2026 · 被引用 1 次
- Exploring Selective Avoidance for Online User Behavior Analysis: A Forest of Thought ExplanationXiaohua Wu, Lin Li, Kaize Shi, Xiaohui Tao 等AAAI 2026
它引用的顶会 Paper7
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent VariableSiqi Bao, Huang He, Fan Wang, Hua Wu 等ACL 2020 · 被引用 229 次
- Controlled Text Generation as Continuous Optimization with Multiple ConstraintsSachin Kumar, Eric Malmi, Aliaksei Severyn, Yulia TsvetkovNeurIPS 2021 · 被引用 91 次
- Generating Counter Narratives against Online Hate Speech: Data and StrategiesSerra Sinem Tekiroglu, Yi-Ling Chung, Marco GueriniACL 2020 · 被引用 13 次
- Counterspeakers' Perspectives: Unveiling Barriers and AI Needs in the Fight against Online HateJimin Mun, Cathy Buerger, Jenny T. Liang, Joshua Garland 等CHI 2024 · 被引用 12 次
相关 Paper
- Is Safer Better? The Impact of Guardrails on the Argumentative Strength of LLMs in Hate Speech CounteringHelena Bonaldi, Greta Damo, Nicolás Benjamín Ocampo, Elena Cabrio 等EMNLP 2024 · 被引用 2 次
- LLM-based Rewriting of Inappropriate Argumentation using Reinforcement Learning from Machine FeedbackTimon Ziegenbein, Gabriella Skitalinskaya, Alireza Bayat Makou, Henning WachsmuthACL 2024
- A Fine-Grained Taxonomy of Replies to Hate SpeechXinchen Yu, Ashley Zhao, Eduardo Blanco, Lingzi HongEMNLP 2023 · 被引用 1 次
- F²RL: Factuality and Faithfulness Reinforcement Learning Framework for Claim-Guided Evidence-Supported Counterspeech GenerationHaiyang Wang, Yuchen Pan, Xin Song, Xuechen Zhao 等EMNLP 2024 · 被引用 1 次
- Comparing human and LLM politeness strategies in free productionHaoran Zhao, Robert D. HawkinsEMNLP 2025 · 被引用 2 次
