"They are uncultured": Unveiling Covert Harms and Social Threats in LLM Generated Conversations
Preetam Prabhu Srikar Dammu, Hayoung Jung, Anjali Singh, Monojit Choudhury, Tanushree Mitra
摘要
Large language models (LLMs) have emerged as an integral part of modern societies, powering user-facing applications such as personal assistants and enterprise applications like recruitment tools. Despite their utility, research indicates that LLMs perpetuate systemic biases. Yet, prior works on LLM harms predominantly focus on Western concepts like race and gender, often overlooking cultural concepts from other parts of the world. Additionally, these studies typically investigate "harm" as a singular dimension, ignoring the various and subtle forms in which harms manifest. To address this gap, we introduce the Covert Harms and Social Threats (CHAST), a set of seven metrics grounded in social science literature. We utilize evaluation models aligned with human assessments to examine the presence of covert harms in LLM-generated conversations, particularly in the context of recruitment. Our experiments reveal that seven out of the eight LLMs included in this study generated conversations riddled with CHAST, characterized by malign views expressed in seemingly neutral language unlikely to be detected by existing methods. Notably, these LLMs manifested more extreme views and opinions when dealing with non-Western concepts like caste, compared to Western ones such as race. Warning: This paper has instances of offensive language to serve as examples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Parasites in the Toolchain: A Large-Scale Analysis of Attacks on the MCP EcosystemShuli Zhao, Qinsheng Hou, Zihan Zhan, Yanhao Wang 等S&P 2026 · 被引用 20 次
- FairI Tales: Evaluation of Fairness in Indian Contexts with a Focus on Bias and StereotypesJanki Atul Nawale, Mohammed Safi Ur Rahman Khan, Janani D, Mansi Gupta 等ACL 2025 · 被引用 5 次
- Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender BiasesBufan Gao, Elisa KreissEMNLP 2025 · 被引用 1 次
- Do LLM Agents Mirror Socio-Cognitive Effects in Power-Asymmetric Conversations?Anvesh Rao Vijjini, Sagar Manjunath, Snigdha ChaturvediACL 2026
- FilBench: Can LLMs Understand and Generate Filipino?Lester James Validad Miranda, Elyanah Aco, Conner G. Manuel, Jan Christian Blaise Cruz 等EMNLP 2025
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Co-Writing Screenplays and Theatre Scripts with Language Models: Evaluation by Industry ProfessionalsPiotr Mirowski, Kory W. Mathewson, Jaylen Pittman, Richard EvansCHI 2023 · 被引用 235 次
- An Empirical Survey of the Effectiveness of Debiasing Techniques for Pre-trained Language ModelsNicholas Meade, Elinor Poole-Dayan, Siva ReddyACL 2022 · 被引用 160 次
相关 Paper
- CoBia: Constructed Conversations Can Trigger Otherwise Concealed Societal Biases in LLMsNafiseh Nikeghbal, Amir Hossein Kargaran, Jana DiesnerEMNLP 2025
- Probing Social Bias in Labor Market Text Generation by ChatGPT: A Masked Language Model ApproachLei Ding, Yang Hu, Nicole Denier, Enze Shi 等NeurIPS 2024 · 被引用 4 次
- Generative AI and Perceptual Harms: Who's Suspected of using LLMs?Kowe Kadoma, Danaë Metaxa, Mor NaamanCHI 2025 · 被引用 28 次
- First-Person Fairness in ChatbotsTyna Eloundou, Alex Beutel, David G. Robinson, Keren Gu 等ICLR 2025 · 被引用 3 次
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 被引用 495 次
