Evaluating LLMs in Open-Source Games
Swadesh Sistla, Max Kleiman-Weiner
摘要
Large Language Models' (LLMs) programming capabilities enable their participation in open-source games: a game-theoretic setting in which players submit computer programs in lieu of actions. These programs offer numerous advantages, including interpretability, inter-agent transparency, and formal verifiability; additionally, they enable program equilibria, solutions that leverage the transparency of code and are inaccessible within normal-form settings. We evaluate the capabilities of leading open-and closed-weight LLMs to predict and classify program strategies and evaluate features of the approximate program equilibria reached by LLM agents in dyadic and evolutionary settings. We identify the emergence of payoffmaximizing, cooperative, and deceptive strategies, characterize the adaptation of mechanisms within these programs over repeated open-source games, and analyze their comparative evolutionary fitness. We find that open-source games serve as a viable environment to study and steer the emergence of cooperative strategy in multi-agent dilemmas.
def strategy(self, opponent) def strategy(self, opponent)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social DilemmasEmanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer 等ICML 2026 · 被引用 15 次
- EIP: Weighted Ranking of LLMs by Quantifying Question DifficultyXingjian Hu, Ziqian Zhang, Yue Huang, Kai Zhang 等ICLR 2026
它引用的顶会 Paper5
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris 等UIST 2023 · 被引用 1,882 次
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 被引用 53 次
- Similarity-based cooperative equilibriumCaspar Oesterheld, Johannes Treutlein, Roger B. Grosse, Vincent Conitzer 等NeurIPS 2023 · 被引用 8 次
- Modeling Others' Minds as CodeKunal Jha, Aydan Yuenan Huang, Eric Ye, Natasha Jaques 等ICLR 2026 · 被引用 6 次
- Cross-environment Cooperation Enables Zero-shot Multi-agent CoordinationKunal Jha, Wilka Carvalho, Yancheng Liang, Simon Shaolei Du 等ICML 2025
相关 Paper
- Discovering Differences in Strategic Behavior between Humans and LLMsCaroline L Wang, Daniel Kasenberg, Kimberly Stachenfeld, Pablo Samuel CastroICML 2026 · 被引用 1 次
- Code World Models for General Game PlayingWolfgang Lehrach, Daniel Hennes, Miguel Lazaro-Gredilla, Xinghua Lou 等ICLR 2026 · 被引用 27 次
- Game-Theoretic Co-Evolution for LLM-Based Heuristic DiscoveryXinyi Ke, Kai Li, Junliang Xing, Yifan Zhang 等ICML 2026 · 被引用 1 次
- ProxyWar: Dynamic Assessment of LLM Code Generation in Game ArenasWenjun Peng, Xinyu Wang, Qi WuICSE 2026
- Opponent Shaping in LLM AgentsMarta Emili Garcia Segura, Stephen Hailes, Mirco MusolesiICLR 2026 · 被引用 3 次
