Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem Solving
Iseo Kim, Eunjin Hong, Juae Kim
摘要
This study investigates how different moral conditions influence the general problem-solving capabilities of Large Language Models (LLMs). We examine whether the role of morality in human decision-making can also serve as a useful lens for analyzing variation in LLM behavior. Specifically, we define distinct moral conditions based on Kohlberg's theory of moral development and design prompts intended to elicit model outputs aligned with each condition. The validity of this alignment is assessed using the Defining Issues Test, a human evaluation tool. We then evaluate models under each condition on the MMLU benchmark, which measures general problem-solving ability across diverse domains. Experimental results show that different moral perspectives correspond to differences in model behavior during general reasoning, as reflected in both responses and internal representations. In particular, more advanced moral conditions tend to elicit more reflective reasoning patterns, which are often linked to improved performance. Our study broadens the scope of LLM morality, which has traditionally been examined mainly in ethical judgment settings. More broadly, it suggests that morality may function not only as a mechanism for safety alignment but also as a factor that shapes model behavior during reasoning. Code and data are available at https: //github.com/ISEOKIM/llm-morality
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 被引用 682 次
- Evaluating the Moral Beliefs Encoded in LLMsNino Scherrer, Claudia Shi, Amir Feder, David M. BleiNeurIPS 2023 · 被引用 316 次
相关 Paper
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human BehaviorsTiancheng Hu, Joachim Baumann, Lorenzo Lupo, Nigel Collier 等ICLR 2026 · 被引用 61 次
- Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-SortsJingting Zheng, Yuqi Ren, Linhao Yu, Yongqi Leng 等ACL 2026
- Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation FrameworkMohna Chakraborty, Lu Wang, David JurgensEMNLP 2025
- MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than OutcomesYu Ying Chiu, Michael S. Lee, Rachel Calcott, Brandon Handoko 等ICLR 2026 · 被引用 23 次
- Deontological Keyword Bias: The Impact of Modal Expressions on Normative Judgments of Language ModelsBumjin Park, Leejinsil Leejinsil, Jaesik ChoiACL 2025
