Hate Personified: Investigating the role of LLMs in content moderation
Sarah Masud, Sahajpreet Singh, Viktor Hangya, Alexander Fraser, Tanmoy Chakraborty
摘要
For subjective tasks such as hate detection, where people perceive hate differently, the Large Language Model's (LLM) ability to represent diverse groups is unclear. By including additional context in prompts, we comprehensively analyze LLM's sensitivity to geographical priming, persona attributes, and numerical information to assess how well the needs of various groups are reflected. Our findings on two LLMs, five languages, and six datasets reveal that mimicking persona-based attributes leads to annotation variability. Meanwhile, incorporating geographical signals leads to better regional alignment. We also find that the LLMs are sensitive to numerical anchors, indicating the ability to leverage community-based flagging efforts and exposure to adversaries. Our work provides preliminary guidelines and highlights the nuances of applying LLMs in culturally sensitive cases. 1 * Equal Contribution 1 Disclaimer: The paper contains examples of strong and hateful language. I like my girlfriends like I like my dogs Rescued from a young age and stays in their cage. "Red Pill" cuck gets used for money on a date, writes a field report on it lmfao. This is how they work. They are domestic terrorists. They are taking over corporations world wide and nothing good will come of it.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- SMARTER: A Data-efficient Framework to Improve Toxicity Detection with Explanation via Self-augmenting Large Language ModelsHuy Nghiem, Advik Sachdeva, Hal Daumé IIIACL 2026 · 被引用 1 次
- PerspectiveMod: A Perspectivist Resource for Deliberative ModerationEva Maria Vecchi, Neele Falk, Carlotta Quensel, Iman Jundi 等EMNLP 2025
- CHAIRO: Contextual Hierarchical Analogical Induction and Reasoning Optimization for LLMsHaotian Lu, Yuchen Mou, Bingzhe WuACL 2026
它引用的顶会 Paper10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM PromptsJ. D. Zamfirescu-Pereira, Richmond Y. Wong, Bjoern Hartmann, Qian YangCHI 2023 · 被引用 892 次
- Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare QueriesYiqiao Jin, Mohit Chandra, Gaurav Verma, Yibo Hu 等WWW 2024 · 被引用 126 次
- ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech DetectionBadr AlKhamissi, Faisal Ladhak, Srini Iyer, Veselin Stoyanov 等EMNLP 2022 · 被引用 17 次
相关 Paper
- Quantifying the Persona Effect in LLM SimulationsTiancheng Hu, Nigel CollierACL 2024 · 被引用 22 次
- Which Demographics do LLMs Default to During Annotation?Johannes Schäfer, Aidan Combs, Christopher Bagdon, Jiahui Li 等ACL 2025 · 被引用 11 次
- Can Third Parties Read Our Emotions?Jiayi Li, Yingfan Zhou, Pranav Narayanan Venkit, Halima Binte Islam 等ACL 2025 · 被引用 5 次
- Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text PerceptionsMatthias Orlikowski, Jiaxin Pei, Paul Röttger, Philipp Cimiano 等ACL 2025 · 被引用 34 次
- There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language ModelsFriedemann Lipphardt, Moonis Ali, Martin Banzer, Anja Feldmann 等NDSS 2026 · 被引用 1 次
