STEER: Assessing the Economic Rationality of Large Language Models
Narun Krishnamurthi Raman, Taylor Lundy, Samuel Joseph Amouyal, Yoav Levine, Kevin Leyton-Brown, Moshe Tennenholtz
摘要
There is increasing interest in using LLMs as decision-making "agents." Doing so includes many degrees of freedom: which model should be used; how should it be prompted; should it be asked to introspect, conduct chain-of-thought reasoning, etc? Settling these questions-and more broadly, determining whether an LLM agent is reliable enough to be trusted-requires a methodology for assessing such an agent's economic rationality. In this paper, we provide one. We begin by surveying the economic literature on rational decision making, taxonomizing a large set of finegrained "elements" that an agent should exhibit, along with dependencies between them. We then propose a benchmark distribution called STEER (Systematic and Tuneable Evaluation of Economic Rationality) that quantitatively scores an LLMs performance on these elements and, combined with a user-provided rubric, produces a "STEER report card." Finally, we describe the results of a large-scale empirical experiment with 14 different LLMs, characterizing the both current state of the art and the impact of different model sizes on models' ability to exhibit rational behavior. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIsMantas Mazeika, Xuwang Yin, Rishub Tamirisa, Jaehyuk Lim 等NeurIPS 2025 · 被引用 84 次
- Is Your LLM Overcharging You? Tokenization, Transparency, and IncentivesAnder Artola Velasco, Stratis Tsirtsis, Nastaran Okati, Manuel Gomez-RodriguezICML 2026 · 被引用 16 次
- Do Large Language Models Know What They Are Capable Of?Casey O. Barkan, Sidney Black, Oliver SourbutICLR 2026 · 被引用 11 次
- Reasoning Models Are Test Exploiters: Rethinking Multiple ChoiceNarun Raman, Taylor Lundy, Kevin Leyton-BrownICML 2026 · 被引用 10 次
- Are Large Language Models Sensitive to the Motives Behind Communication?Addison J. Wu, Ryan Liu, Kerem Oktar, Theodore R. Sumers 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper13
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu 等NeurIPS 2023 · 被引用 5,989 次
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
相关 Paper
- Distributive Fairness in Large Language Models: Evaluating Alignment with Human ValuesHadi Hosseini, Samarth KhannaNeurIPS 2025 · 被引用 14 次
- SteerConf: Steering LLMs for Confidence ElicitationZiang Zhou, Tianyuan Jin, Jieming Shi, Qing LiNeurIPS 2025 · 被引用 23 次
- STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language ModelsKai Chen, Zihao He, Taiwei Shi, Kristina LermanEMNLP 2025 · 被引用 1 次
- Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-MakingYuanjun Feng, Vivek Choudhary, Yash Raj ShresthaEMNLP 2025 · 被引用 2 次
- Market-Bench: Benchmarking Large Language Models on Economic and Trade CompetitionYushuo Zheng, Huiyu Duan, Zicheng Zhang, Yucheng Zhu 等ACL 2026 · 被引用 1 次
