Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech Detection
Min Zhang, Jianfeng He, Taoran Ji, Chang-Tien Lu
摘要
The fairness and trustworthiness of Large Language Models (LLMs) are receiving increasing attention. Implicit hate speech, which employs indirect language to convey hateful intentions, occupies a significant portion of practice. However, the extent to which LLMs effectively address this issue remains insufficiently examined. This paper delves into the capability of LLMs to detect implicit hate speech (Classification Task) and express confidence in their responses (Calibration Task). Our evaluation meticulously considers various prompt patterns and mainstream uncertainty estimation methods. Our findings highlight that LLMs exhibit two extremes: (1) LLMs display excessive sensitivity towards groups or topics that may cause fairness issues, resulting in misclassifying benign statements as hate speech. (2) LLMs' confidence scores for each method excessively concentrate on a fixed range, remaining unchanged regardless of the dataset's complexity. Consequently, the calibration performance is heavily reliant on primary classification accuracy. These discoveries unveil new limitations of LLMs, underscoring the need for caution when optimizing models to ensure they do not veer towards extremes. This serves as a reminder to carefully consider sensitivity and confidence in the pursuit of model fairness.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Influences on LLM Calibration: A Study of Response Agreement, Loss Functions, and Prompt StylesYuxi Xia, Pedro Henrique Luz de Araujo, Klim Zaporojets, Benjamin RothACL 2025 · 被引用 12 次
- On the Risk of Evidence Pollution for Malicious Social Text Detection in the Era of LLMsHerun Wan, Minnan Luo, Zhixiong Su, Guang Dai 等ACL 2025 · 被引用 5 次
- Making FETCH! Happen: Finding Emergent Dog Whistles Through Common HabitatsKuleen Sasse, Carlos Alejandro Aguirre, Isabel Cachola, Sharon Levy 等ACL 2025 · 被引用 3 次
- Evaluating Large Language Models for Detecting AntisemitismJay Patel, Hrudayangam Mehta, Jeremy BlackburnEMNLP 2025 · 被引用 1 次
- MetaFaith: Faithful Natural Language Uncertainty Expression in LLMsGabrielle Kaili-May Liu, Gal Yona, Avi Caciularu, Idan Szpektor 等EMNLP 2025
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Can LLMs Express Their Uncertainty? An Empirical Evaluation of Confidence Elicitation in LLMsMiao Xiong, Zhiyuan Hu, Xinyang Lu, Yifei Li 等ICLR 2024 · 被引用 867 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Mix-n-Match : Ensemble and Compositional Methods for Uncertainty Calibration in Deep LearningJize Zhang, Bhavya Kailkhura, Thomas Yong-Jin HanICML 2020 · 被引用 276 次
相关 Paper
- Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech DetectionSanjeevan Selvaganapathy, Mehwish NasimACL 2026
- Large Language Models Must Be Taught to Know What They Don't KnowSanyam Kapoor, Nate Gruver, Manley Roberts, Katie Collins 等NeurIPS 2024 · 被引用 124 次
- How to Correctly Report LLM-as-a-Judge EvaluationsChungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn 等ICML 2026 · 被引用 24 次
- Sheep's Skin, Wolf's Deeds: Are LLMs Ready for Metaphorical Implicit Hate Speech?Jingjie Zeng, Liang Yang, Zekun Wang, Yuanyuan Sun 等ACL 2025 · 被引用 4 次
- Rethinking Implicit Hate Speech Detection: Focusing on Latent Hate Components via Dual-Process ArgumentationShiqi Sun, Du Su, Wei Chen, Xueqi ChengWWW 2026
