Many LLMs Are More Utilitarian Than One
Anita Keshmirian, Razan Baltaji, Babak Hemmatian, Hadi Asghari, Lav R. Varshney
摘要
Moral judgment is integral to large language models' (LLMs) social reasoning. As multi-agent systems gain prominence, it becomes crucial to understand how LLMs function when collaborating compared to operating as individual agents. In human moral judgment, group deliberation leads to a Utilitarian Boost: a tendency to endorse norm violations that inflict harm but maximize benefits for the greatest number of people. We study whether a similar dynamic emerges in multi-agent LLM systems. We test six models on well-established sets of moral dilemmas across two conditions: (1) Solo, where models reason independently, and (2) Group, where they engage in multi-turn discussions in pairs or triads. In personal dilemmas, where agents decide whether to directly harm an individual for the benefit of others, all models rated moral violations as more acceptable when part of a group, demonstrating a Utilitarian Boost similar to that observed in humans. However, the mechanism for the boost in LLMs differed: While humans in groups become more utilitarian due to heightened sensitivity to decision outcomes, LLM groups showed diverse profiles, for example, reduced sensitivity to norms or enhanced impartiality. We report model differences in when and how strongly the boost manifests. We also discuss prompt and agent compositions that enhance or mitigate the effect. We end with a discussion of the implications for AI alignment, multi-agent design, and artificial moral reasoning. Code available at: https://github.com/baltaci-r/MoralAgents Recent work draws on insights from social psychology to investigate emergent distortions in grouplevel multi-agent LLM reasoning, demonstrating phenomena such as conformity [10,11], belief †
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Evaluating the Moral Beliefs Encoded in LLMsNino Scherrer, Claudia Shi, Amir Feder, David M. BleiNeurIPS 2023 · 被引用 316 次
- Secret Collusion among AI Agents: Multi-Agent Deception via SteganographySumeet Ramesh Motwani, Mikhail Baranchuk, Martin Strohmeier, Vijay Bolina 等NeurIPS 2024 · 被引用 140 次
- Modular Pluralism: Pluralistic Alignment via Multi-LLM CollaborationShangbin Feng, Taylor Sorensen, Yuhan Liu, Jillian Fisher 等EMNLP 2024 · 被引用 12 次
- Exploring Collaboration Mechanisms for LLM Agents: A Social Psychology ViewJintian Zhang, Xin Xu, Ningyu Zhang, Ruibo Liu 等ACL 2024
- Do as We Do, Not as You Think: the Conformity of Large Language ModelsZhiyuan Weng, Guikun Chen, Wenguan WangICLR 2025
相关 Paper
- Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM CollectivesChanggeon Ko, Jisu Shin, Hoyun Song, Huije Lee 等ACL 2026 · 被引用 1 次
- Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem SolvingIseo Kim, Eunjin Hong, Juae KimACL 2026
- From Single to Societal: Analyzing Persona-Induced Bias in Multi-Agent InteractionsJiayi Li, Xiao Liu, Yansong FengAAAI 2026 · 被引用 3 次
- Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion DynamicsJiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng 等EMNLP 2025
- Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social InteractionsMaojia Song, Pala Tej Deep, Ruiwen Zhou, Weisheng Jin 等ICLR 2026
