Sheep's Skin, Wolf's Deeds: Are LLMs Ready for Metaphorical Implicit Hate Speech?
Jingjie Zeng, Liang Yang, Zekun Wang, Yuanyuan Sun, Hongfei Lin
摘要
Implicit hate speech has become a significant challenge for online platforms, as it often avoids detection by large language models (LLMs) due to its indirectly expressed hateful intent. This study identifies the limitations of LLMs in detecting implicit hate speech, particularly when disguised as seemingly harmless expressions in a rhetorical device. To address this challenge, we employ a Jailbreaking strategy and Energy-based Constrained Decoding techniques, and design a small model for measuring the energy of metaphorical rhetoric. This approach can lead to LLMs generating metaphorical implicit hate speech. Our research reveals that advanced LLMs, like GPT-4o, frequently misinterpret metaphorical implicit hate speech, and fail to prevent its propagation effectively. Even specialized models, like ShieldGemma and LlamaGuard, demonstrate inadequacies in blocking such content, often misclassifying it as harmless speech. This work points out the vulnerability of current LLMs to implicit hate speech, and emphasizes the improvements to address hate speech threats better.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMsShiyao Cui, Qinglin Zhang, Di Wang, Yida Lu 等ACL 2026
- When in Doubt, Consult: Expert Debate for Sexism Detection via Confidence-Based RoutingAnwar Alajmi, Gabriele PergolaACL 2026
- Reinforcement Learning-Guided Adaptive Tuning for Out-of-Distribution Harmful Text DetectionMengyu Xiang, Tinghao Chen, Boxu Han, Qiudan Li 等ACL 2026
- SoftHateBench: Evaluating Moderation Models Against Reasoning-Driven, Policy-Compliant HostilityXuanyu Su, Diana Inkpen, Nathalie JapkowiczWWW 2026
它引用的顶会 Paper8
- Jailbroken: How Does LLM Safety Training Fail?Alexander Wei, Nika Haghtalab, Jacob SteinhardtNeurIPS 2023 · 被引用 2,230 次
- Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak SupervisionCollin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker 等ICML 2024 · 被引用 443 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- COLD Decoding: Energy-based Constrained Text Generation with Langevin DynamicsLianhui Qin, Sean Welleck, Daniel Khashabi, Yejin ChoiNeurIPS 2022 · 被引用 217 次
- COLD-Attack: Jailbreaking LLMs with Stealthiness and ControllabilityXingang Guo, Fangxu Yu, Huan Zhang, Lianhui Qin 等ICML 2024 · 被引用 173 次
相关 Paper
- from Benign import Toxic: Jailbreaking the Language Model via Adversarial MetaphorsYu Yan, Sheng Sun, Zenghao Duan, Teli Liu 等ACL 2025 · 被引用 14 次
- Emoji Attack: Enhancing Jailbreak Attacks Against Judge LLM DetectionZhipeng Wei, Yuqi Liu, N. Benjamin ErichsonICML 2025
- URLcoat: Exploiting Web Search Capability to Jailbreak Large Language ModelsYiheng Sun, Linkang Du, Zhou Su, Yuntao Wang 等S&P 2026
- Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech DetectionMin Zhang, Jianfeng He, Taoran Ji, Chang-Tien LuACL 2024
- GraphShield: Graph-Theoretic Modeling of Network-Level Dynamics for Robust Jailbreak DetectionSunghee Dong, Sungwon Yi, Kangmin Bae, Jaeyoon Kim 等ICLR 2026
