PREDICT: Multi-Agent-based Debate Simulation for Generalized Hate Speech Detection
Someen Park, Jaehoon Kim, Seungwan Jin, Sohyun Park, Kyungsik Han
摘要
While a few public benchmarks have been proposed for training hate speech detection models, the differences in labeling criteria between these benchmarks pose challenges for generalized learning, limiting the applicability of the models.Previous research has presented methods to generalize models through data integration or augmentation, but overcoming the differences in labeling criteria between datasets remains a limitation.To address these challenges, we propose PREDICT, a novel framework that uses the notion of multi-agent for hate speech detection.PREDICT consists of two phases: (1) PRE (Perspectivebased REasoning): Multiple agents are created based on the induced labeling criteria of given datasets, and each agent generates stances and reasons; (2) DICT (Debate using InCongruenT references): Agents representing hate and nonhate stances conduct the debate, and a judge agent classifies hate or non-hate and provides a balanced reason.Experiments on five representative public benchmarks show that PREDICT achieves superior cross-evaluation performance compared to methods that focus on specific labeling criteria or majority voting methods.Furthermore, we validate that PREDICT effectively mediates differences between agents' opinions and appropriately incorporates minority opinions to reach a consensus.Our code is available at https://github.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language ModelsChen Han, Wenzhen Zheng, Xijin TangEMNLP 2025 · 被引用 2 次
- RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech DetectionYejin Lee, Hyeseon An, Yo-Sub HanACL 2026
- Beyond Single-View Detection: A Dual-Space Reasoning Framework for Interpretable Harmful Meme UnderstandingWenqing Hou, Hongkui Tu, Ye Wang, Yue Zhang 等ACL 2026
- SeMob: Semantic Synthesis for Dynamic Urban Mobility PredictionRunfei Chen, Shuyang Jiang, Wei HuangEMNLP 2025
- Tracing Belief-Driven Thoughts with Theory-of-Mind Agents: An Opinion Analysis FrameworkJintao Wen, Yunfeng Ning, Hankun Kang, Xin Miao 等WWW 2026
它引用的顶会 Paper6
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum 等ICML 2024 · 被引用 1,562 次
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent DebateChi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu 等ICLR 2024 · 被引用 871 次
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent DebateTian Liang, Zhiwei He, Wenxiang Jiao, Xing Wang 等EMNLP 2024 · 被引用 177 次
- KOLD: Korean Offensive Language DatasetYounghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn 等EMNLP 2022 · 被引用 41 次
- Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language ModelsAbhishek Kumar, Sarfaroz Yunusov, Ali EmamiACL 2024 · 被引用 3 次
相关 Paper
- Hate Speech Detection Based on Sentiment Knowledge SharingXianbing Zhou, Yang Yong, Xiaochao Fan, Ge Ren 等ACL 2021
- SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant HostilityXuanyu Su, Diana Inkpen, Nathalie JapkowiczWWW 2026
- Spanning the Spectrum of Hatred Detection: A Persian Multi-Label Hate Speech Dataset with Annotator RationalesZahra Delbari, Nafise Sadat Moosavi, Mohammad Taher PilehvarAAAI 2024 · 被引用 11 次
- When the Majority is Wrong: Modeling Annotator Disagreement for Subjective TasksEve Fleisig, Rediet Abebe, Dan KleinEMNLP 2023 · 被引用 11 次
- MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance DetectionWeihai Lu, Zhejun Zhao, Yanshu Li, Huan HeACL 2026
