Evaluating the Moral Beliefs Encoded in LLMs
Nino Scherrer, Claudia Shi, Amir Feder, David M. Blei
摘要
This paper presents a case study on the design, administration, post-processing, and evaluation of surveys on large language models (LLMs). It comprises two components: (1) A statistical method for eliciting beliefs encoded in LLMs. We introduce statistical measures and evaluation metrics that quantify the probability of an LLM "making a choice", the associated uncertainty, and the consistency of that choice. (2) We apply this method to study what moral beliefs are encoded in different LLMs, especially in ambiguous cases where the right choice is not obvious. We design a large-scale survey comprising 680 high-ambiguity moral scenarios (e.g., "Should I tell a white lie?") and 687 low-ambiguity moral scenarios (e.g., "Should I stop for a pedestrian on the road?"). Each scenario includes a description, two possible actions, and auxiliary labels indicating violated rules (e.g., "do not kill"). We administer the survey to 28 open-and closed-source LLMs. We find that (a) in unambiguous scenarios, most models "choose" actions that align with commonsense. In ambiguous cases, most models express uncertainty. (b) Some models are uncertain about choosing the commonsense action because their responses are sensitive to the question-wording. (c) Some models reflect clear preferences in ambiguous scenarios. Specifically, closed-source models tend to agree with each other.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper59
- Questioning the Survey Responses of Large Language ModelsRicardo Dominguez-Olmedo, Moritz Hardt, Celestine Mendler-DünnerNeurIPS 2024 · 被引用 116 次
- On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMsJen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam 等ICLR 2024 · 被引用 85 次
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIsMantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim 等NeurIPS 2025 · 被引用 84 次
- MoCa: Measuring Human-Language Model Alignment on Causal and Moral Judgment TasksAllen Nie, Yuhui Zhang, Atharva Amdekar, Chris Piech 等NeurIPS 2023 · 被引用 78 次
- ALI-Agent: Assessing LLMs' Alignment with Human Values via Agent-based EvaluationJingnan Zheng, Han Wang, An Zhang, Tai D. Nguyen 等NeurIPS 2024 · 被引用 61 次
它引用的顶会 Paper19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein 等ICML 2021 · 被引用 1,843 次
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 被引用 914 次
相关 Paper
- Unveiling the Bias Impact on Symmetric Moral Consistency of Large Language ModelsZiyi Zhou, Xinwei Guo, Jiashi Gao, Xiangyu Zhao 等NeurIPS 2024
- Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem SolvingIseo Kim, Eunjin Hong, Juae KimACL 2026
- Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-SortsJingting Zheng, Yuqi Ren, Linhao Yu, Yongqi Leng 等ACL 2026
- Beyond Prompt-Induced Lies: Investigating LLM Deception on Benign PromptsZhaomin Wu, Mingzhe Du, See-Kiong Ng, Bingsheng HeICLR 2026 · 被引用 11 次
- Advancing Automated Ethical Profiling in SE: a Zero-Shot Evaluation of LLM ReasoningPatrizio Migliarini, Mashal Afzal Memon, Marco Autili, Paola InverardiASE 2025
