PREDICT: Multi-Agent-based Debate Simulation for Generalized Hate Speech Detection
Someen Park, Jaehoon Kim, Seungwan Jin, Sohyun Park, Kyungsik Han
Abstract
While a few public benchmarks have been proposed for training hate speech detection models, the differences in labeling criteria between these benchmarks pose challenges for generalized learning, limiting the applicability of the models.Previous research has presented methods to generalize models through data integration or augmentation, but overcoming the differences in labeling criteria between datasets remains a limitation.To address these challenges, we propose PREDICT, a novel framework that uses the notion of multi-agent for hate speech detection.PREDICT consists of two phases: (1) PRE (Perspectivebased REasoning): Multiple agents are created based on the induced labeling criteria of given datasets, and each agent generates stances and reasons; (2) DICT (Debate using InCongruenT references): Agents representing hate and nonhate stances conduct the debate, and a judge agent classifies hate or non-hate and provides a balanced reason.Experiments on five representative public benchmarks show that PREDICT achieves superior cross-evaluation performance compared to methods that focus on specific labeling criteria or majority voting methods.Furthermore, we validate that PREDICT effectively mediates differences between agents' opinions and appropriately incorporates minority opinions to reach a consensus.Our code is available at https://github.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cbbbfe00-6349-4683-8695-adee4f1ab465Cited by top-tier papers7
- Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language ModelsChen Han, Wenzhen Zheng, Xijin TangEMNLP 2025 · 2 citations
- RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech DetectionYejin Lee, Hyeseon An, Yo-Sub HanACL 2026
- Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme UnderstandingWenqing Hou, Hongkui Tu, Ye Wang, Yue Zhang et al.ACL 2026
- SeMob: Semantic Synthesis for Dynamic Urban Mobility PredictionRunfei Chen, Shuyang Jiang, Wei HuangEMNLP 2025
- Tracing Belief-Driven Thoughts with Theory-of-Mind Agents: An Opinion Analysis FrameworkJintao Wen, Yunfeng Ning, Hankun Kang, Xin Miao et al.WWW 2026
Builds on6
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu et al.ICLR 2024 · 871 citations
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateTian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang et al.EMNLP 2024 · 177 citations
- KOLD: Korean Offensive Language DatasetYounghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn et al.EMNLP 2022 · 41 citations
- Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language ModelsAbhishek Kumar, Sarfaroz Yunusov, Ali EmamiACL 2024 · 3 citations
Related papers
- Hate Speech Detection Based on Sentiment Knowledge SharingXianbing Zhou, Yang Yong, Xiaochao Fan, Ge Ren et al.ACL 2021
- SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant HostilityXuanyu Su, Diana Inkpen, Nathalie JapkowiczWWW 2026
- Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator RationalesZahra Delbari, Nafise Sadat Moosavi, Mohammad Taher PilehvarAAAI 2024 · 11 citations
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 11 citations
- MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance DetectionWeihai Lu, Zhejun Zhao, Yanshu Li, Huan HeACL 2026
