Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context
Jingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara, Deming Chen
Abstract
When making decisions under uncertainty, individuals often deviate from rational behavior, which can be evaluated across three dimensions: risk preference, probability weighting, and loss aversion. Given the widespread use of large language models (LLMs) in decision-making processes, it is crucial to assess whether their behavior aligns with human norms and ethical expectations or exhibits potential biases. Several empirical studies have investigated the rationality and social behavior performance of LLMs, yet their internal decision-making tendencies and capabilities remain inadequately understood. This paper proposes a framework, grounded in behavioral economics, to evaluate the decision-making behaviors of LLMs. Through a multiple-choice-list experiment, we estimate the degree of risk preference, probability weighting, and loss aversion in a context-free setting for three commercial LLMs: ChatGPT-4.0-Turbo, Claude-3-Opus, and Gemini-1.0-pro. Our results reveal that LLMs generally exhibit patterns similar to humans, such as risk aversion and loss aversion, with a tendency to overweight small probabilities. However, there are significant variations in the degree to which these behaviors are expressed across different LLMs. We also explore their behavior when embedded with socio-demographic features, uncovering significant disparities. For instance, when modeled with attributes of sexual minority groups or physical disabilities, Claude-3-Opus displays increased risk aversion, leading to more conservative choices. These findings underscore the need for careful consideration of the ethical implications and potential biases in deploying LLMs in decision-making scenarios. Therefore, this study advocates for developing standards and guidelines to ensure that LLMs operate within ethical boundaries while enhancing their utility in complex decision-making environments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4aa0ea4b-0ad5-49f1-ac91-d4e2300a45baCited by top-tier papers12
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- LLM Strategic Reasoning: Agentic Study through Behavioral Game TheoryJingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara et al.NeurIPS 2025 · 23 citations
- Constrained Discrete DiffusionMichael Cardei, Jacob K. Christopher, Bhavya Kailkhura, Tom Hartvigsen et al.NeurIPS 2025 · 20 citations
- Do Large Language Models Know What They Are Capable Of?Casey O. Barkan, Sidney Black, Oliver SourbutICLR 2026 · 11 citations
- Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-MakingYuanjun Feng, Vivek Choudhary, Yash Raj ShresthaEMNLP 2025 · 2 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Large Language Models are Human-Level Prompt EngineersYongchao Zhou, Andrei Ioan Muresanu, Ziwen Han, Keiran Paster et al.ICLR 2023 · 297 citations
Related papers
- Evaluating and Aligning Human Economic Risk Preferences in LLMsJiaxin Liu, Yixuan Tang, Yi Yang, Kar Yan TamEMNLP 2025
- To Mask or to Mirror: Human-AI Alignment in Collective ReasoningCrystal Qian, Aaron T. Parisi, Clémentine Bouleau, Vivian Tsai et al.EMNLP 2025 · 1 citation
- Distributive Fairness in Large Language Models: Evaluating Alignment with Human ValuesHadi Hosseini, Samarth KhannaNeurIPS 2025 · 14 citations
- Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMsLuise Ge, Yongyan Zhang, Yevgeniy VorobeychikACL 2026
- Large Language Models Assume People are More Rational than We Really areRyan Liu, Jiayi Geng, Joshua C. Peterson, Ilia Sucholutsky et al.ICLR 2025
