Multi-turn Evaluation of Anthropomorphic Behaviours in Large Language Models
Lujain Ibrahim, Canfer Akbulut, Rasmi Elasmar, Charvi Rastogi, Minsuk Kahng, Meredith Ringel Morris, Kevin R. McKee, Verena Rieser, Murray Shanahan, Laura Weidinger
摘要
The tendency of users to anthropomorphise large language models (LLMs) is of growing societal interest. Here, we present AnthroBench, a novel empirical method and tool 1 for evaluating anthropomorphic LLM behaviours in realistic settings. Our work introduces three key advances; first, we develop a multi-turn evaluation of 14 distinct anthropomorphic behaviours, moving beyond single-turn assessment. Second, we present a scalable, automated approach by leveraging simulations of user interactions, enabling efficient and reproducible assessment. Third, we conduct an interactive, large-scale human subject study (N = 1101) to empirically validate that the model behaviours we measure predict real users' anthropomorphic perceptions. We find that all evaluated LLMs exhibit similar behaviours, primarily characterised by relationship-building (e.g., empathy and validation) with users and first-person pronoun use. Crucially, we observe that the majority of these anthropomorphic behaviours only first occur after multiple turns, underscoring the necessity of multi-turn evaluations for understanding complex social phenomena in human-AI interaction. Our work provides a robust empirical foundation for investigating how design choices influence anthropomorphic model behaviours and for progressing the ethical debate on the desirability of these behaviours.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMsPengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath 等ICML 2026 · 被引用 33 次
- Generative Value Conflicts Reveal LLM PrioritiesAndy Liu, Kshitish Ghate, Mona T. Diab, Daniel Fried 等ICLR 2026 · 被引用 17 次
- Thinking beyond the anthropomorphic paradigm benefits LLM researchLujain Ibrahim, Myra ChengACL 2026 · 被引用 13 次
- Humanizing Machines: Rethinking LLM Anthropomorphism Through a Multi-Level Framework of DesignYunze Xiao, Lynnette Hui Xian Ng, Jiarui Liu, Mona T. DiabEMNLP 2025 · 被引用 2 次
- Evaluating LLMs When They Do Not Know the Answer: Statistical Evaluation of Mathematical Reasoning via Comparative SignalsZihan Dong, Zhixian Zhang, Yang Zhou, Can Jin 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LLM Evaluators Recognize and Favor Their Own GenerationsArjun Panickssery, Samuel R. Bowman, Shi FengNeurIPS 2024 · 被引用 865 次
- WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language ModelsLiwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger 等NeurIPS 2024 · 被引用 247 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Mirages. On Anthropomorphism in Dialogue SystemsGavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser 等EMNLP 2023 · 被引用 44 次
相关 Paper
- DarkBench: Benchmarking Dark Patterns in Large Language ModelsEsben Kran, Jord Nguyen, Akash Kundu, Sami Jawhar 等ICLR 2025
- ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language ModelsYuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang 等ACL 2024
- Therapy as an NLP Task: Comparing LLMs and Human Peers Behaviors in CBT SessionsZainab Iftikhar, Sean Ransom, Amy Wei Xiao, Nicole Nugent 等CSCW 2026
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human BehaviorsTiancheng Hu, Joachim Baumann, Lorenzo Lupo, Nigel Collier 等ICLR 2026 · 被引用 61 次
- Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with HumansJen-tse Huang, Man Ho Lam, Eric John Li, Shujie Ren 等NeurIPS 2024 · 被引用 63 次
