Mind the Value-Action Gap: Do LLMs Act in Alignment with Their Values?
Hua Shen, Nicholas Clark, Tanu Mitra
摘要
Existing research primarily evaluates the values of LLMs by examining their stated inclinations towards specific values. However, the"Value-Action Gap,"a phenomenon rooted in environmental and social psychology, reveals discrepancies between individuals'stated values and their actions in real-world contexts. To what extent do LLMs exhibit a similar gap between their stated values and their actions informed by those values? This study introduces ValueActionLens, an evaluation framework to assess the alignment between LLMs'stated values and their value-informed actions. The framework encompasses the generation of a dataset comprising 14.8k value-informed actions across twelve cultures and eleven social topics, and two tasks to evaluate how well LLMs'stated value inclinations and value-informed actions align across three different alignment measures. Extensive experiments reveal that the alignment between LLMs'stated values and actions is sub-optimal, varying significantly across scenarios and models. Analysis of misaligned results identifies potential harms from certain value-action gaps. To predict the value-action gaps, we also uncover that leveraging reasoned explanations improves performance. These findings underscore the risks of relying solely on the LLMs'stated values to predict their behaviors and emphasize the importance of context-aware evaluations of LLM values and value-action gaps.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- SimBench: Benchmarking the Ability of Large Language Models to Simulate Human BehaviorsTiancheng Hu, Joachim Baumann, Lorenzo Lupo, Nigel Collier 等ICLR 2026 · 被引用 61 次
- Distributive Fairness in Large Language Models: Evaluating Alignment with Human ValuesHadi Hosseini, Samarth KhannaNeurIPS 2025 · 被引用 14 次
- The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language ModelsYike Shi, Qing Xiao, Qing Hu, Hong Shen 等CHI 2026 · 被引用 7 次
- Deep Value Benchmark: Measuring Whether Models Generalize Deep Values or Shallow PreferencesJoshua Ashkinaze, Hua Shen, Sai Avula, Eric Gilbert 等NeurIPS 2025 · 被引用 7 次
- Whose Facts Win? LLM Source Preferences under Knowledge ConflictsJakob Schuster, Vagrant Gautam, Katja MarkertACL 2026 · 被引用 3 次
它引用的顶会 Paper7
- Human-AI Interaction in Human Resource Management: Understanding Why Employees Resist Algorithmic Evaluation at Workplaces and How to Mitigate BurdensHyanghee Park, Daehwan Ahn, Kartik Hosanagar, Joonhwan LeeCHI 2021 · 被引用 117 次
- A Framework of Severity for Harmful Content OnlineMorgan Klaus Scheuerman, Jialun Aaron Jiang, Casey Fiesler, Jed R. BrubakerCSCW 2021 · 被引用 113 次
- Designing Responsible AI: Adaptations of UX Practice to Meet Responsible AI ChallengesQiaosi Wang, Michael Madaio, Shaun K. Kane, Shivani Kapania 等CHI 2023 · 被引用 90 次
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman 等ACL 2020 · 被引用 36 次
- Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language ModelsPaul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck 等ACL 2024
相关 Paper
- Unintended Harms of Value-Aligned LLMs: Psychological and Empirical InsightsSooyung Choi, Jaehyeok Lee, Xiaoyuan Yi, Jing Yao 等ACL 2025
- ValueBench: Towards Comprehensively Evaluating Value Orientations and Understanding of Large Language ModelsYuanyi Ren, Haoran Ye, Hanjun Fang, Xin Zhang 等ACL 2024
- Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid ItemsJongwook Han, Dongmin Choi, Woojung Song, Eun-Ju Lee 等ACL 2025
- Expectation Alignment of Language Models for Real-World User ExpectationsMiaomiao Li, Yang Wang, Bin Liang, Shudong Liu 等ICML 2026
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language ModelsRuilin Yao, Bo Zhang, Jirui Huang, Xinwei Long 等ICLR 2026 · 被引用 8 次
