Tractable Multi-Agent Reinforcement Learning through Behavioral Economics
Eric Mazumdar, Kishan Panaganti, Laixi Shi
Abstract
A significant roadblock to the development of principled multi-agent reinforcement learning (MARL) algorithms is the fact that desired solution concepts like Nash equilibria may be intractable to compute. We show how one can overcome this obstacle by introducing concepts from behavioral economics into MARL. To do so, we imbue agents with two key features of human decision-making: risk aversion and bounded rationality. We show that introducing these two properties into games gives rise to a class of equilibria-risk-averse quantal response equilibria (RQE)-which are tractable to compute in all n-player matrix and finite-horizon Markov games. In particular, we show that they emerge as the endpoint of noregret learning in suitably adjusted versions of the games. Crucially, the class of computationally tractable RQE is independent of the underlying game structure and only depends on agents' degrees of risk-aversion and bounded rationality. To validate the expressivity of this class of solution concepts we show that it captures peoples' patterns of play in a number of 2-player matrix games previously studied in experimental economics. Furthermore, we give a first analysis of the sample complexity of computing these equilibria in finite-horizon Markov games when one has access to a generative model. We validate our findings on a simple multiagent reinforcement learning benchmark. Our results open the doors for to the development of new decentralized multi-agent reinforcement learning algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 16446704-be37-490e-9684-36a403fa92adCited by top-tier papers1
Ask how each one uses itBuilds on19
- Nash Learning from Human FeedbackRémi Munos, Michal Valko, Daniele Calandriello, Mohammad Gheshlaghi Azar et al.ICML 2024 · 212 citations
- Near-Optimal Reinforcement Learning with Self-PlayYu Bai, Chi Jin, Tiancheng YuNeurIPS 2020 · 150 citations
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 130 citations
- Robust Multi-Agent Reinforcement Learning with Model UncertaintyKaiqing Zhang, Tao Sun, Yunzhe Tao, Sahika Genc et al.NeurIPS 2020 · 118 citations
- Fast Policy Extragradient Methods for Competitive Games with Entropy RegularizationShicong Cen, Yuting Wei, Yuejie ChiNeurIPS 2021 · 105 citations
Related papers
- Breaking the Curse of Multiagency in Robust Multi-Agent Reinforcement LearningLaixi Shi, Jingchu Gai, Eric Mazumdar, Yuejie Chi et al.ICML 2025
- Bounded Rationality Equilibrium Learning in Mean Field GamesYannick Eich, Christian Fabian, Kai Cui, Heinz KoepplAAAI 2025 · 2 citations
- From Behavioral Theories to Econometrics: Inferring Preferences of Human Agents from Data on Repeated InteractionsGali NotiAAAI 2021 · 7 citations
- Robust Adversarial Reinforcement Learning via Bounded Rationality CurriculaAryaman Reddi, Maximilian Tölle, Jan Peters, Georgia Chalvatzaki et al.ICLR 2024 · 11 citations
- Exploration-Exploitation in Multi-Agent Competition: Convergence with Bounded RationalityStefanos Leonardos, Georgios Piliouras, Kelly SpendloveNeurIPS 2021 · 43 citations
