Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation Framework
Mohna Chakraborty, Lu Wang, David Jurgens
摘要
Large language models (LLMs) are increasingly deployed in domains requiring moral understanding, yet their reasoning often remains shallow, and misaligned with human reasoning (Jiang et al., 2021). Unlike humans, whose moral reasoning integrates contextual trade-offs, value systems, and ethical theories, LLMs often rely on surface patterns, leading to biased decisions in morally and ethically complex scenarios. To address this gap, we present a value-grounded framework for evaluating and distilling structured moral reasoning in LLMs. We benchmark 12 open-source models across four moral datasets using a taxonomy of prompts grounded in value systems, ethical theories, and cognitive reasoning strategies. Our evaluation is guided by four questions: (1) Does reasoning improve LLM decision-making over direct prompting? (2) Which types of value/ethical frameworks most effectively guide LLM reasoning? (3) Which cognitive reasoning strategies lead to better moral performance? (4) Can small-sized LLMs acquire moral competence through distillation? We find that prompting with explicit moral structure consistently improves accuracy and coherence, with first-principles reasoning and Schwartz's + care-ethics scaffolds yielding the strongest gains. Furthermore, our supervised distillation approach transfers moral competence from large to small models without additional inference cost. Together, our results offer a scalable path toward interpretable and value-grounded models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple PerspectivesAyoung Lee, Ryan Sungmo Kwon, Peter Railton, Lu WangICLR 2026 · 被引用 13 次
- Pluralistic Alignment for Healthcare: A Role-Driven FrameworkJiayou Zhong, Anudeex Shetty, Chao Jia, Xuanrui Lin 等EMNLP 2025 · 被引用 1 次
- Where Paths Split: Localized, Calibrated Control of Moral Reasoning in Large Language ModelsChenchen Yuan, Zheyu Zhang, Gjergji KasneciACL 2026
- Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion DynamicsJiarui Liu, Yueqi Song, Yunze Xiao, Mingqian Zheng 等EMNLP 2025
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
- Self-Alignment with Instruction BacktranslationXian Li, Ping Yu, Chunting Zhou, Timo Schick 等ICLR 2024 · 被引用 174 次
- Reasoning with Language Model is Planning with World ModelShibo Hao, Yi Gu, Haodi Ma, Joshua Jiahua Hong 等EMNLP 2023 · 被引用 109 次
相关 Paper
- Beyond Value Benchmarks: Measuring Value-Structure Alignment in Large Language Models via Symmetric Q-SortsJingting Zheng, Yuqi Ren, Linhao Yu, Yongqi Leng 等ACL 2026
- SOLAR: Towards Characterizing Subjectivity of Individuals through Modeling Value Conflicts and Trade-offsYounghun Lee, Dan GoldwasserEMNLP 2025
- Reasoning Scaffolding: Distilling the Flow of Thought from LLMsXiangyu Wen, Junhua Huang, Zeju Li, Min Li 等ICLR 2026 · 被引用 7 次
- Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem SolvingIseo Kim, Eunjin Hong, Juae KimACL 2026
- Advancing Automated Ethical Profiling in SE: a Zero-Shot Evaluation of LLM ReasoningPatrizio Migliarini, Mashal Afzal Memon, Marco Autili, Paola InverardiASE 2025
