LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory
Jingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara, Deming Chen
Abstract
What does it truly mean for a language model to "reason" strategically, and can scaling up alone guarantee intelligent, context-aware decisions? Strategic decisionmaking requires adaptive reasoning, where agents anticipate and respond to others' actions under uncertainty. Yet, most evaluations of large language models (LLMs) for strategic decision-making often rely heavily on Nash Equilibrium (NE) benchmarks, overlook reasoning depth, and fail to reveal the mechanisms behind model behavior. To address this gap, we introduce a behavioral game-theoretic evaluation framework that disentangles intrinsic reasoning from contextual influence. Using this framework, we evaluate 22 state-of-the-art LLMs across diverse strategic scenarios. We find models like GPT-o3-mini, GPT-o1, and DeepSeek-R1 lead in reasoning depth. Through thinking chain analysis, we identify distinct reasoning styles-such as maximin or belief-based strategies-and show that longer reasoning chains do not consistently yield better decisions. Furthermore, embedding demographic personas reveals context-sensitive shifts: some models (e.g., GPT-4o, Claude-3-Opus) improve when assigned female identities, while others (e.g., Gemini 2.0) show diminished reasoning under minority sexuality personas. These findings underscore that technical sophistication alone is insufficient; alignment with ethical standards, human expectations, and situational nuance is essential for the responsible deployment of LLMs in interactive settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fbabf68c-2614-48ac-be0c-e8816d674037Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 651 citations
- Evaluating the Moral Beliefs Encoded in LLMsNino Scherrer, Claudia Shi, Amir Feder, David M. BleiNeurIPS 2023 · 316 citations
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsShashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan et al.ICLR 2024 · 212 citations
- Can Large Language Models Serve as Rational Players in Game Theory? A Systematic AnalysisCaoyun Fan, Jindou Chen, Yaohui Jin, Hao HeAAAI 2024 · 123 citations
- Decision-Making Behavior Evaluation Framework for LLMs under Uncertain ContextJingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara et al.NeurIPS 2024 · 71 citations
Related papers
- To Mask or to Mirror: Human-AI Alignment in Collective ReasoningCrystal Qian, Aaron T. Parisi, Clémentine Bouleau, Vivian Tsai et al.EMNLP 2025 · 1 citation
- GTBench: Uncovering the Strategic Reasoning Capabilities of LLMs via Game-Theoretic EvaluationsJinhao Duan, Renming Zhang, James Diffenderfer, Bhavya Kailkhura et al.NeurIPS 2024 · 79 citations
- Competing Large Language Models in Multi-Agent Gaming EnvironmentsJen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang et al.ICLR 2025
- Reasoning Models Are Test Exploiters: Rethinking Multiple ChoiceNarun Raman, Taylor Lundy, Kevin Leyton-BrownICML 2026 · 10 citations
- InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning StylesZizhen Li, Chuanhao Li, Yibin Wang, Qi Chen et al.EMNLP 2025
