Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning
Aswini Kumar Padhi, Anil Bandhakavi, Tanmoy Chakraborty
摘要
Counterspeech has proven to be a powerful tool to combat hate speech online. Previous studies have focused on generating counterspeech conditioned only on specific strategies (single attributed). However, a holistic approach considering multiple attributes simultaneously can yield more nuanced and effective responses. Here, we introduce HiPPrO, Hierarchical Prefix learning with Preference Optimization, a novel two-stage framework that utilizes the effectiveness of attribute-specific prefix embedding spaces hierarchically optimized during the counterspeech generation process in the first phase. Thereafter, we incorporate both reference and reward-free preference optimization to generate more constructive counterspeech. Furthermore, we extend IntentCONANv2 by annotating all 13, 973 counterspeech instances with emotion labels by five annotators. HiPPrO leverages hierarchical prefix optimization to integrate these dual attributes effectively. An extensive evaluation demonstrates that HiPPrO achieves a ∼ 38% improvement in strategy conformity and a ∼ 3%, ∼ 2%, ∼ 3% improvement in Rouge-1, Rouge-2, and Rouge-L, respectively, compared to several baseline models. Human evaluations further substantiate the superiority of our approach, highlighting the enhanced relevance and appropriateness of the generated counterspeech. This work underscores the potential of multiattribute conditioning in advancing the efficacy of counterspeech generation systems. 1 Our code is available on Github and dataset is opensourced on Hugging-face.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 被引用 94 次
- Hate is the New Infodemic: A Topic-aware Modeling of Hate Speech Diffusion on TwitterSarah Masud, Subhabrata Dutta, Sakshi Makkar, Chhavi Jain 等ICDE 2021 · 被引用 39 次
- Human-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech CounteringHelena Bonaldi, Sara Dellantonio, Serra Sinem Tekiroglu, Marco GueriniEMNLP 2022 · 被引用 18 次
相关 Paper
- Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech GenerationRishabh Gupta, Shaily Desai, Manvi Goel, Anil Bandhakavi 等ACL 2023 · 被引用 10 次
- Outcome-Constrained Large Language Models for Countering Hate SpeechLingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying SongEMNLP 2024 · 被引用 5 次
- F²RL: Factuality and Faithfulness Reinforcement Learning Framework for Claim-Guided Evidence-Supported Counterspeech GenerationHaiyang Wang, Yuchen Pan, Xin Song, Xuechen Zhao 等EMNLP 2024 · 被引用 1 次
- HABERTOR: An Efficient and Effective Deep Hatespeech DetectorThanh Tran, Yifan Hu, Changwei Hu, Kevin Yen 等EMNLP 2020
- PREDICT: Multi-Agent-based Debate Simulation for Generalized Hate Speech DetectionSomeen Park, Jaehoon Kim, Seungwan Jin, Sohyun Park 等EMNLP 2024 · 被引用 5 次
