Do LLM Agents Have Regret? A Case Study in Online Learning and Games
Chanwoo Park, Xiangyu Liu, Asuman E. Ozdaglar, Kaiqing Zhang
Abstract
Large language models (LLMs) have been increasingly employed for (interactive) decisionmaking, via the development of LLM-based autonomous agents. Despite their emerging successes, the performance of LLM agents in decision-making has not been fully investigated through quantitative metrics-especially in the multi-agent setting when they interact with each other, a typical scenario in real-world LLM-agent applications. To better understand the limits of LLM agents in these interactive environments, we propose to study their interactions in benchmark decision-making settings in online learning and game theory, through the performance metric of regret. We first empirically study the no-regret behaviors of LLMs in canonical non-stochastic online learning problems, as well as the emergence of equilibria when multiple of them interact through playing repeated games. We then provide some theoretical insights into sublinear regret growth in the cases we observed, under certain assumptions on (supervised) pre-training and the data generation model. Notably, we also identify (simple) cases where advanced LLMs such as GPT-4 fail to be no-regret. To further promote the no-regret behaviors, we propose a novel unsupervised training loss, the regret-loss, which, in contrast to the supervised pre-training loss, does not require the labels of (optimal) actions. Finally, we establish the statistical guarantee of generalization bound for regret-loss minimization, and more importantly, the optimization guarantee that minimizing such a loss can lead to known no-regret learning algorithms, when single-layer self-attention models are used. Our further experiments demonstrate the effectiveness of our regret-loss, especially in addressing the above "regrettable" cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4b64b63-2797-4560-87a8-fd6b308fd744Cited by top-tier papers13
- MDAgents: An Adaptive Collaboration of LLMs for Medical Decision-MakingYubin Kim, Chanwoo Park, Hyewon Jeong, Yik Siu Chan et al.NeurIPS 2024 · 291 citations
- Can large language models explore in-context?Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster, Cyril Zhang et al.NeurIPS 2024 · 95 citations
- MAPoRL: Multi-Agent Post-Co-Training for Collaborative Large Language Models with Reinforcement LearningChanwoo Park, Seungju Han, Xingzhi Guo, Asuman E. Ozdaglar et al.ACL 2025 · 63 citations
- Reward Is Enough: LLMs Are In-Context Reinforcement LearnersKefan Song, Amir Moeini, Peng Wang, Lei Gong et al.ICLR 2026 · 42 citations
- Large Language Models Think Too Fast To Explore EffectivelyLan Pan, Hanbo Xie, Robert C. WilsonNeurIPS 2025 · 10 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- PaLM-E: An Embodied Multimodal Language ModelDanny Driess, Fei Xia, Mehdi S. M. Sajjadi, Corey Lynch et al.ICML 2023 · 2,601 citations
Related papers
- Post-Training LLMs as Better Decision-Making Agents: A Regret-Minimization ApproachChanwoo Park, Ziyang Chen, Asuman Ozdaglar, Kaiqing ZhangICML 2026 · 5 citations
- Competing Large Language Models in Multi-Agent Gaming EnvironmentsJen-tse Huang, Eric John Li, Man Ho Lam, Tian Liang et al.ICLR 2025
- When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM TrainingSanxing Chen, Xiaoyin Chen, Yukun Huang, Roy Xie et al.ICLR 2026 · 3 citations
- Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the Machiavelli BenchmarkAlexander Pan, Jun Shern Chan, Andy Zou, Nathaniel Li et al.ICML 2023 · 200 citations
- Benchmarking Large Language Model Capabilities for Conditional GenerationJoshua Maynez, Priyanka Agrawal, Sebastian GehrmannACL 2023 · 6 citations
