Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning
Aswini Kumar Padhi, Anil Bandhakavi, Tanmoy Chakraborty
Abstract
Counterspeech has proven to be a powerful tool to combat hate speech online. Previous studies have focused on generating counterspeech conditioned only on specific strategies (single attributed). However, a holistic approach considering multiple attributes simultaneously can yield more nuanced and effective responses. Here, we introduce HiPPrO, Hierarchical Prefix learning with Preference Optimization, a novel two-stage framework that utilizes the effectiveness of attribute-specific prefix embedding spaces hierarchically optimized during the counterspeech generation process in the first phase. Thereafter, we incorporate both reference and reward-free preference optimization to generate more constructive counterspeech. Furthermore, we extend IntentCONANv2 by annotating all 13, 973 counterspeech instances with emotion labels by five annotators. HiPPrO leverages hierarchical prefix optimization to integrate these dual attributes effectively. An extensive evaluation demonstrates that HiPPrO achieves a ∼ 38% improvement in strategy conformity and a ∼ 3%, ∼ 2%, ∼ 3% improvement in Rouge-1, Rouge-2, and Rouge-L, respectively, compared to several baseline models. Human evaluations further substantiate the superiority of our approach, highlighting the enhanced relevance and appropriateness of the generated counterspeech. This work underscores the potential of multiattribute conditioning in advancing the efficacy of counterspeech generation systems. 1 Our code is available on Github and dataset is opensourced on Hugging-face.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d3bef261-3a4f-4fbe-b286-35481a14ef07Builds on8
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- The Power of Scale for Parameter-Efficient Prompt TuningBrian Lester, Rami Al-Rfou, Noah ConstantEMNLP 2021 · 94 citations
- Hate is the New Infodemic: A Topic-aware Modeling of Hate Speech Diffusion on TwitterSarah Masud, Subhabrata Dutta, Sakshi Makkar, Chhavi Jain et al.ICDE 2021 · 39 citations
- Human-Machine Collaboration Approaches to Build a Dialogue Dataset for Hate Speech CounteringHelena Bonaldi, Sara Dellantonio, Serra Sinem Tekiroglu, Marco GueriniEMNLP 2022 · 18 citations
Related papers
- Counterspeeches up my sleeve! Intent Distribution Learning and Persistent Fusion for Intent-Conditioned Counterspeech GenerationRishabh Gupta, Shaily Desai, Manvi Goel, Anil Bandhakavi et al.ACL 2023 · 10 citations
- Outcome-Constrained Large Language Models for Countering Hate SpeechLingzi Hong, Pengcheng Luo, Eduardo Blanco, Xiaoying SongEMNLP 2024 · 5 citations
- F²RL: Factuality and Faithfulness Reinforcement Learning Framework for Claim-Guided Evidence-Supported Counterspeech GenerationHaiyang Wang, Yuchen Pan, Xin Song, Xuechen Zhao et al.EMNLP 2024 · 1 citation
- HABERTOR: An Efficient and Effective Deep Hatespeech DetectorThanh Tran, Yifan Hu, Changwei Hu, Kevin Yen et al.EMNLP 2020
- PREDICT: Multi-Agent-based Debate Simulation for Generalized Hate Speech DetectionSomeen Park, Jaehoon Kim, Seungwan Jin, Sohyun Park et al.EMNLP 2024 · 5 citations
