Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral
Shivani Kumar, David Jurgens
Abstract
Moral reasoning is a complex cognitive process shaped by individual experiences and cultural contexts and presents unique challenges for computational analysis. While natural language processing (NLP) offers promising tools for studying this phenomenon, current research lacks cohesion, employing discordant datasets and tasks that examine isolated aspects of moral reasoning. We bridge this gap with UNI-MORAL, a unified dataset integrating psychologically grounded and social-media-derived moral dilemmas annotated with labels for action choices, ethical principles, contributing factors, and consequences, alongside annotators' moral and cultural profiles. Recognizing the cultural relativity of moral reasoning, UNI-MORAL spans six languages, Arabic, Chinese, English, Hindi, Russian, and Spanish, capturing diverse socio-cultural contexts. We demonstrate UNIMORAL's utility through a benchmark evaluations of three large language models (LLMs) across four tasks: action prediction, moral typology classification, factor attribution analysis, and consequence generation. Key findings reveal that while implicitly embedded moral contexts enhance the moral reasoning capability of LLMs, there remains a critical need for increasingly specialized approaches to further advance moral reasoning in these models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e58c33b5-d9e2-456a-8fc0-809ba2aecc50Cited by top-tier papers2
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic PerspectivesYinuo Xu, Veronica Derricks, Allison Earl, David JurgensACL 2026 · 8 citations
- Structured Moral Reasoning in Language Models: A Value-Grounded Evaluation FrameworkMohna Chakraborty, Lu Wang, David JurgensEMNLP 2025
Builds on7
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch et al.ICLR 2021 · 878 citations
- SCRUPLES: A Corpus of Community Ethical Judgments on 32, 000 Real-Life AnecdotesNicholas Lourie, Ronan Le Bras, Yejin ChoiAAAI 2021 · 147 citations
- When to Make Exceptions: Exploring Language Models as Accounts of Human Moral JudgmentZhijing Jin, Sydney Levine, Fernando Gonzalez Adauto, Ojasv Kamal et al.NeurIPS 2022 · 146 citations
- Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their ConsequencesDenis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes et al.EMNLP 2021 · 83 citations
Related papers
- UniM: A Unified Any-to-Any Interleaved Multimodal BenchmarkYanlin Li, Minghui Guo, Kaiwen Zhang, Shize Zhang et al.CVPR 2026 · 10 citations
- Do Morals Guide How LLMs Think? The Role of Ethical Perspectives in General Problem SolvingIseo Kim, Eunjin Hong, Juae KimACL 2026
- UniRRM: Unified Reasoning Reward Models Across Languages and Evaluation ParadigmsPeng Lai, Yichao Du, Junchao Wu, Weibo Gao et al.ICML 2026
- HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language ModelsZhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia et al.ICLR 2026 · 24 citations
- MMAC: A Multilingual, Multimodal Alignment Framework for Cultural Grounding EvaluationWeihua Zheng, Zhengyuan Liu, Tanmoy Chakraborty, Weiwen Xu et al.ACL 2026
