SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant Hostility
Xuanyu Su, Diana Inkpen, Nathalie Japkowicz
摘要
Online hate on social media ranges from overt slurs and threats (hard hate speech) to soft hate speech: discourse that appears reasonable on the surface but uses framing and value-based arguments to steer audiences toward blaming or excluding a target group. We hypothesize that current moderation systems, largely optimized for surface toxicity cues, are not robust to this reasoning-driven hostility, yet existing benchmarks do not measure this gap systematically. We introduce SoftHateBench, a generative benchmark that produces soft-hate variants while preserving the underlying hostile standpoint. To generate soft hate, we integrate the Argumentum Model of Topics (AMT) and Relevance Theory (RT) in a unified framework: AMT provides the backbone argument structure for rewriting an explicit hateful standpoint into a seemingly neutral discussion while preserving the stance, and RT guides generation to keep the AMT chain logically coherent. The benchmark spans 7 sociocultural domains and 28 target groups, comprising 4,745 soft-hate instances. Evaluations across encoder-based detectors, general-purpose LLMs, and safety models show a consistent drop from hard to soft tiers: systems that detect explicit hostility often fail when the same stance is conveyed through subtle, reasoningbased language. Disclaimer. Contains offensive examples used solely for research. CCS Concepts • Computing methodologies → Artificial intelligence; Natural language processing; Natural language generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
- CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic NetworkSreyan Ghosh, Manan Suri, Purva Chiniya, Utkarsh Tyagi 等EMNLP 2023 · 被引用 9 次
- Sheep's Skin, Wolf's Deeds: Are LLMs Ready for Metaphorical Implicit Hate Speech?Jingjie Zeng, Liang Yang, Zekun Wang, Yuanyuan Sun 等ACL 2025 · 被引用 4 次
- Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe KielaACL 2021
- HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate CampaignsXinyue Shen, Yixin Wu, Yiting Qu, Michael Backes 等USENIX Security 2025
相关 Paper
- PREDICT: Multi-Agent-based Debate Simulation for Generalized Hate Speech DetectionSomeen Park, Jaehoon Kim, Seungwan Jin, Sohyun Park 等EMNLP 2024 · 被引用 5 次
- Rethinking Implicit Hate Speech Detection: Focusing on Latent Hate Components via Dual-Process ArgumentationShiqi Sun, Du Su, Wei Chen, Xueqi ChengWWW 2026
- AmpleHate: Amplifying the Attention for Versatile Implicit Hate DetectionYejin Lee, Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub HanEMNLP 2025 · 被引用 2 次
- Drifting Away from Truth: GenAI-Driven News Diversity Challenges LVLM-Based Misinformation DetectionFanxiao Li, Jiaying Wu, Tingchao Fu, Yunyun Dong 等AAAI 2026 · 被引用 4 次
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius 等KDD 2024 · 被引用 7 次
