Evaluating LLMs in Open-Source Games
Swadesh Sistla, Max Kleiman-Weiner
Abstract
Large Language Models' (LLMs) programming capabilities enable their participation in open-source games: a game-theoretic setting in which players submit computer programs in lieu of actions. These programs offer numerous advantages, including interpretability, inter-agent transparency, and formal verifiability; additionally, they enable program equilibria, solutions that leverage the transparency of code and are inaccessible within normal-form settings. We evaluate the capabilities of leading open-and closed-weight LLMs to predict and classify program strategies and evaluate features of the approximate program equilibria reached by LLM agents in dyadic and evolutionary settings. We identify the emergence of payoffmaximizing, cooperative, and deceptive strategies, characterize the adaptation of mechanisms within these programs over repeated open-source games, and analyze their comparative evolutionary fitness. We find that open-source games serve as a viable environment to study and steer the emergence of cooperative strategy in multi-agent dilemmas.
def strategy(self, opponent) def strategy(self, opponent)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 87bde9bd-1687-425c-a7bf-131421b61d57Cited by top-tier papers2
- CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social DilemmasEmanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer et al.ICML 2026 · 15 citations
- EIP: Weighted Ranking of LLMs by Quantifying Question DifficultyXingjian Hu, Ziqian Zhang, Yue Huang, Kai Zhang et al.ICLR 2026
Builds on5
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 53 citations
- Similarity-based cooperative equilibriumCaspar Oesterheld, Johannes Treutlein, Roger B. Grosse, Vincent Conitzer et al.NeurIPS 2023 · 8 citations
- Modeling Others' Minds as CodeKunal Jha, Aydan Yuenan Huang, Eric Ye, Natasha Jaques et al.ICLR 2026 · 6 citations
- Cross-environment Cooperation Enables Zero-shot Multi-agent CoordinationKunal Jha, Wilka Carvalho, Yancheng Liang, Simon Shaolei Du et al.ICML 2025
Related papers
- Discovering Differences in Strategic Behavior between Humans and LLMsCaroline L Wang, Daniel Kasenberg, Kimberly Stachenfeld, Pablo Samuel CastroICML 2026 · 1 citation
- Code World Models for General Game PlayingWolfgang Lehrach, Daniel Hennes, Miguel Lazaro-Gredilla, Xinghua Lou et al.ICLR 2026 · 27 citations
- Game-Theoretic Co-Evolution for LLM-Based Heuristic DiscoveryXinyi Ke, Kai Li, Junliang Xing, Yifan Zhang et al.ICML 2026 · 1 citation
- ProxyWar: Dynamic Assessment of LLM Code Generation in Game ArenasWenjun Peng, Xinyu Wang, Qi WuICSE 2026
- Opponent Shaping in LLM AgentsMarta Emili Garcia Segura, Stephen Hailes, Mirco MusolesiICLR 2026 · 3 citations
