SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant Hostility
Xuanyu Su, Diana Inkpen, Nathalie Japkowicz
Abstract
Online hate on social media ranges from overt slurs and threats (hard hate speech) to soft hate speech: discourse that appears reasonable on the surface but uses framing and value-based arguments to steer audiences toward blaming or excluding a target group. We hypothesize that current moderation systems, largely optimized for surface toxicity cues, are not robust to this reasoning-driven hostility, yet existing benchmarks do not measure this gap systematically. We introduce SoftHateBench, a generative benchmark that produces soft-hate variants while preserving the underlying hostile standpoint. To generate soft hate, we integrate the Argumentum Model of Topics (AMT) and Relevance Theory (RT) in a unified framework: AMT provides the backbone argument structure for rewriting an explicit hateful standpoint into a seemingly neutral discussion while preserving the stance, and RT guides generation to keep the AMT chain logically coherent. The benchmark spans 7 sociocultural domains and 28 target groups, comprising 4,745 soft-hate instances. Evaluations across encoder-based detectors, general-purpose LLMs, and safety models show a consistent drop from hard to soft tiers: systems that detect explicit hostility often fail when the same stance is conveyed through subtle, reasoningbased language. Disclaimer. Contains offensive examples used solely for research. CCS Concepts • Computing methodologies → Artificial intelligence; Natural language processing; Natural language generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fdfa5b52-6b50-4a27-bbb7-101fbdce915dBuilds on9
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi et al.EMNLP 2021 · 159 citations
- CoSyn: Detecting Implicit Hate Speech in Online Conversations Using a Context Synergized Hyperbolic NetworkSreyan Ghosh, Manan Suri, Purva Chiniya, Utkarsh Tyagi et al.EMNLP 2023 · 9 citations
- Sheep's Skin, Wolf's Deeds: Are LLMs Ready for Metaphorical Implicit Hate Speech?Jingjie Zeng, Liang Yang, Zekun Wang, Yuanyuan Sun et al.ACL 2025 · 4 citations
- Learning from the Worst: Dynamically Generated Datasets to Improve Online Hate DetectionBertie Vidgen, Tristan Thrush, Zeerak Waseem, Douwe KielaACL 2021
- HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate CampaignsXinyue Shen, Yixin Wu, Yiting Qu, Michael Backes et al.USENIX Security 2025
Related papers
- PREDICT: Multi-Agent-based Debate Simulation for Generalized Hate Speech DetectionSomeen Park, Jaehoon Kim, Seungwan Jin, Sohyun Park et al.EMNLP 2024 · 5 citations
- Rethinking Implicit Hate Speech Detection: Focusing on Latent Hate Components via Dual-Process ArgumentationShiqi Sun, Du Su, Wei Chen, Xueqi ChengWWW 2026
- AmpleHate: Amplifying the Attention for Versatile Implicit Hate DetectionYejin Lee, Joonghyuk Hahn, Hyeseon Ahn, Yo-Sub HanEMNLP 2025 · 2 citations
- Drifting Away from Truth: GenAI-Driven News Diversity Challenges LVLM-Based Misinformation DetectionFanxiao Li, Jiaying Wu, Tingchao Fu, Yunyun Dong et al.AAAI 2026 · 4 citations
- Hate Speech Detection with Generalizable Target-aware FairnessTong Chen, Danny Wang, Xurong Liang, Marten Risius et al.KDD 2024 · 7 citations
