Lune

ICLR2025

Competing Large Language Models in Multi-Agent Gaming Environments

Jen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang, Wenxuan Wang, Youliang Yuan, Wenxiang Jiao, Xing Wang, Zhaopeng Tu, Michael R. Lyu

2025年份

摘要

➢How is LLMs' decision-making ability in game theoretic scenes? 1. Multiparty: theory-of-mind reasoning 2. Calculation: arithmetic reasoning 3. Understanding: environment & game rules ➢Games: ideal test bed for LLM evaluation 1. Scope: abstraction of real-world scenarios 2. Quantifiability: compute scores with math models 3. Variability: changing game parameters