Confident, Calibrated, or Complicit: Safety Alignment and Ideological Bias in LLM Hate Speech Detection
Sanjeevan Selvaganapathy, Mehwish Nasim
摘要
We investigate the efficacy of Large Language Models (LLMs) in detecting implicit and explicit hate speech, examining how models with minimal safety alignment (uncensored) compare with more heavily aligned (censored) counterparts in a deployed-model setting when deployed using political personas. While uncensored models are often framed as offering a less constrained perspective, our results reveal a trade-off: censored models outperform their uncensored counterparts in both accuracy and robustness, achieving 69.0% versus 64.1% strict accuracy. However, this higher performance is also associated with greater resistance to persona-based influence, while uncensored models are more malleable to ideological framing. Furthermore, we identify critical failures across all models in understanding nuanced language such as irony. We also find alarming fairness disparities in performance across different targeted groups and systemic overconfidence that renders self-reported certainty unreliable. These findings challenge the notion of LLMs as objective arbiters and highlight the need for more sophisticated auditing frameworks that account for fairness, calibration, and ideological consistency. Taken together, these results point to censorship-as-deployed rather than safety alignment in isolation as the more appropriate frame for interpreting model differences.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee 等ICML 2023 · 被引用 764 次
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
- Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language ModelsMyra Cheng, Esin Durmus, Dan JurafskyACL 2023 · 被引用 89 次
相关 Paper
- Don't Go To Extremes: Revealing the Excessive Sensitivity and Calibration Limitations of LLMs in Implicit Hate Speech DetectionMin Zhang, Jianfeng He, Taoran Ji, Chang-Tien LuACL 2024
- Propaganda AI: An Analysis of Semantic Divergence in Large Language ModelsNay Myat Min, Long H. Pham, Yige Li, Jun SunICLR 2026 · 被引用 3 次
- Inertia in Moral and Value Judgments of Large Language ModelsBruce W. Lee, Yeongheon Lee, Hyunsoo ChoACL 2026 · 被引用 5 次
- Too Nice to Tell the Truth: Quantifying Agreeableness-Driven Sycophancy in Role-Playing Language ModelsArya Shah, Deepali Mishra, Chaklam SilpasuwanchaiACL 2026 · 被引用 1 次
- From Pretraining Data to Language Models to Downstream Tasks: Tracking the Trails of Political Biases Leading to Unfair NLP ModelsShangbin Feng, Chan Young Park, Yuhan Liu, Yulia TsvetkovACL 2023 · 被引用 117 次
