HumT DumT: Measuring and controlling human-like language in LLMs
Myra Cheng, Sunny Yu, Dan Jurafsky
摘要
Should LLMs generate language that makes them seem human? Human-like language might improve user experience, but might also lead to deception, overreliance, and stereotyping. Assessing these potential impacts requires a systematic way to measure human-like tone in LLM outputs. We introduce HUMT and SO-CIOT, metrics for human-like tone and other dimensions of social perceptions in text data based on relative probabilities from an LLM. By measuring HUMT across preference and usage datasets, we find that users prefer less human-like outputs from LLMs in many contexts. HUMT also offers insights into the perceptions and impacts of anthropomorphism: human-like LLM outputs are highly correlated with warmth, social closeness, femininity, and low status, which are closely linked to the aforementioned harms. We introduce DUMT, a method using HUMT to systematically control and reduce the degree of human-like tone while preserving model performance. DUMT offers a practical approach for mitigating risks associated with anthropomorphic language generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper14
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 被引用 337 次
- ULTRAFEEDBACK: Boosting Language Models with Scaled AI FeedbackGanqu Cui, Lifan Yuan, Ning Ding, Guanming Yao 等ICML 2024 · 被引用 286 次
- Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language ModelsMyra Cheng, Esin Durmus, Dan JurafskyACL 2023 · 被引用 89 次
- The Illusion of Empathy? Notes on Displays of Emotion in Human-Computer InteractionAndrea Cuadra, Maria Wang, Lynn Andrea Stein, Malte F. Jung 等CHI 2024 · 被引用 75 次
相关 Paper
- Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language ModelsLujain Ibrahim, Canfer Akbulut, Rasmi Elasmar, Charvi Rastogi 等ICLR 2026 · 被引用 40 次
- A Taxonomy of Linguistic Expressions That Contribute To Anthropomorphism of Language TechnologiesAlicia DeVrio, Myra Cheng, Lisa Egede, Alexandra Olteanu 等CHI 2025 · 被引用 28 次
- Dehumanizing Machines: Mitigating Anthropomorphic Behaviors in Text Generation SystemsMyra Cheng, Su Lin Blodgett, Alicia DeVrio, Lisa Egede 等ACL 2025
- Mirages. On Anthropomorphism in Dialogue SystemsGavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser 等EMNLP 2023 · 被引用 44 次
- Comparing human and LLM politeness strategies in free productionHaoran Zhao, Robert D. HawkinsEMNLP 2025 · 被引用 2 次
