MoCa: Measuring Human-Language Model Alignment on Causal and Moral Judgment Tasks
Allen Nie, Yuhui Zhang, Atharva Amdekar, Chris Piech, Tatsunori B. Hashimoto, Tobias Gerstenberg
摘要
Human commonsense understanding of the physical and social world is organized around intuitive theories. These theories support making causal and moral judgments. When something bad happens, we naturally ask: who did what, and why? A rich literature in cognitive science has studied people's causal and moral intuitions. This work has revealed a number of factors that systematically influence people's judgments, such as the violation of norms and whether the harm is avoidable or inevitable. We collected a dataset of stories from 24 cognitive science papers and developed a system to annotate each story with the factors they investigated. Using this dataset, we test whether large language models (LLMs) make causal and moral judgments about text-based scenarios that align with those of human participants. On the aggregate level, alignment has improved with more recent LLMs. However, using statistical analyses, we find that LLMs weigh the different factors quite differently from human participants. These results show how curated, challenge datasets combined with insights from cognitive science can help us go beyond comparisons based merely on aggregate metrics: we uncover LLMs implicit tendencies and show to what extent these align with human intuitions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Aligning to Thousands of Preferences via System Message GeneralizationSeongyun Lee, Sue Hyun Park, Seungone Kim, Minjoon SeoNeurIPS 2024 · 被引用 102 次
- CLASH: Evaluating Language Models on Judging High-Stakes Dilemmas from Multiple PerspectivesAyoung Lee, Ryan Sungmo Kwon, Peter Railton, Lu WangICLR 2026 · 被引用 13 次
- J4R: Learning to Judge with Equivalent Initial State Group Relative Policy OptimizationAustin Xu, Yilun Zhou, Xuan-Phi Nguyen, Caiming Xiong 等ACL 2026 · 被引用 8 次
- Boosting Resilience of Large Language Models through Causality-Driven Robust OptimizationXiaoling Zhou, Mingjie Zhang, Zhemg Lee, Yuncheng Hua 等NeurIPS 2025 · 被引用 5 次
- Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert SelectionTianyi Niu, Justin Chih-Yao Chen, Genta Indra Winata, Shi-Xiong Zhang 等ACL 2026 · 被引用 1 次
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- Aligning AI With Shared Human ValuesDan Hendrycks, Collin Burns, Steven Basart, Andrew Critch 等ICLR 2021 · 被引用 878 次
相关 Paper
- Can Large Language Models Infer Causation from Correlation?Zhijing Jin, Jiarui Liu, Zhiheng Lyu, Spencer Poff 等ICLR 2024 · 被引用 186 次
- RULEBREAKERS: Challenging LLMs at the Crossroads between Formal Logic and Human-like ReasoningJason Chan, Robert J. Gaizauskas, Zhixue ZhaoICML 2025
- SafeText: A Benchmark for Exploring Physical Safety in Language ModelsSharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton 等EMNLP 2022 · 被引用 14 次
- Reading Books is Great, But Not if You Are Driving! Visually Grounded Reasoning about Defeasible Commonsense NormsSeungju Han, Junhyeok Kim, Jack Hessel, Liwei Jiang 等EMNLP 2023
- Can Large Language Models Infer Causal Relationships from Real-World Text?Ryan Saklad, Aman Chadha, Oleg V. Pavlov, Raha MoraffahACL 2026 · 被引用 4 次
