Robust Multi-Agent Reinforcement Learning with Model Uncertainty
Kaiqing Zhang, Tao Sun, Yunzhe Tao, Sahika Genc, Sunil Mallya, Tamer Basar
Abstract
In this work, we study the problem of multi-agent reinforcement learning (MARL) with model uncertainty, which is referred to as robust MARL. This is naturally motivated by some multi-agent applications where each agent may not have perfectly accurate knowledge of the model, e.g., all the reward functions of other agents. Little a priori work on MARL has accounted for such uncertainties, neither in problem formulation nor in algorithm design. In contrast, we model the problem as a robust Markov game, where the goal of all agents is to find policies such that no agent has the incentive to deviate, i.e., reach some equilibrium point, which is also robust to the possible uncertainty of the MARL model. We first introduce the solution concept of robust Nash equilibrium in our setting, and develop a Qlearning algorithm to find such equilibrium policies, with convergence guarantees under certain conditions. In order to handle possibly enormous state-action spaces in practice, we then derive the policy gradients for robust MARL, and develop an actor-critic algorithm with function approximation. Our experiments demonstrate that the proposed algorithm outperforms several baseline MARL methods that do not account for the model uncertainty, in several standard but uncertain cooperative and competitive MARL environments. Equal Contribution 34th Conference on Neural Information Processing Systems (NeurIPS 2020), Vancouver, Canada. * ,t i∈N and π * ,t = N j=1 π j * ,t denote the equilibrium policies of the nature and the equilibrium joint policies of all agents, respectively, computed from Qi t i∈N at time t. The term π 0,i * ,t (s)[a] denotes the a-th element of the policy output π 0,i * ,t (s), a real vector that lies in Ri s ⊆ R |A| . Convergence. Note that convergence of the update (3.3) is in general hard to establish, as the Bellman operator induced by solving a general-sum game in (3.2) does not always satisfy the conditions for the convergence of Q-learning in MDPs and generalized MDPs [42]. As recognized in [25, 26, 27] , convergence of Q-learning in general-sum Markov games indeed requires more conditions. We will establish the convergence of (3.3) under certain conditions, mostly motivated from [25] . Due to space limitation, we defer the results in Supplementary §A.2. The results, though not generally apply to all robust Markov games, provide some proof-of-concept justifications and sanity-check for the convergence of the value-based/Q-learning update. Indeed, developing provable convergent Success Rate
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70d9be2b-70d6-4908-b14e-a2a1bb5dfb76Cited by top-tier papers24
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 157 citations
- LDSA: Learning Dynamic Subtask Assignment in Cooperative Multi-Agent Reinforcement LearningMingyu Yang, Jian Zhao, Xunhan Hu, Wengang Zhou et al.NeurIPS 2022 · 61 citations
- Robust Multi-Agent Coordination via Evolutionary Generation of Auxiliary Adversarial AttackersLei Yuan, Ziqian Zhang, Ke Xue, Hao Yin et al.AAAI 2023 · 31 citations
- Byzantine Robust Cooperative Multi-Agent Reinforcement Learning as a Bayesian GameSimin Li, Jun Guo, Jingqiao Xiu, Ruixiao Xu et al.ICLR 2024 · 30 citations
- Sample-Efficient Robust Multi-Agent Reinforcement Learning in the Face of Environmental UncertaintyLaixi Shi, Eric Mazumdar, Yuejie Chi, Adam WiermanICML 2024 · 23 citations
Builds on2
- Model-Based Multi-Agent RL in Zero-Sum Markov Games with Near-Optimal Sample ComplexityKaiqing Zhang, Sham M. Kakade, Tamer Basar, Lin F. YangNeurIPS 2020 · 144 citations
- Robust Reinforcement Learning for Continuous Control with Model MisspecificationDaniel J. Mankowitz, Nir Levine, Rae Jeong, Abbas Abdolmaleki et al.ICLR 2020 · 138 citations
Related papers
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke et al.NeurIPS 2023 · 17 citations
- Independent Policy Gradient for Large-Scale Markov Potential Games: Sharper Rates, Function Approximation, and Game-Agnostic ConvergenceDongsheng Ding, Chen-Yu Wei, Kaiqing Zhang, Mihailo R. JovanovicICML 2022 · 84 citations
- Sample-Efficient Multi-Agent RL: An Optimization PerspectiveNuoya Xiong, Zhihan Liu, Zhaoran Wang, Zhuoran YangICLR 2024 · 2 citations
- Incentivize without Bonus: Provably Efficient Model-based Online Multi-agent RL for Markov GamesTong Yang, Bo Dai, Lin Xiao, Yuejie ChiICML 2025
- Sample-Efficient Tabular Self-Play for Offline Robust Reinforcement LearningNa Li, Zewu Zheng, Wei Ni, Hangguan Shan et al.NeurIPS 2025 · 1 citation
