Recursive Reasoning Graph for Multi-Agent Reinforcement Learning
Xiaobai Ma, David Isele, Jayesh K. Gupta, Kikuo Fujimura, Mykel J. Kochenderfer
Abstract
Multi-agent reinforcement learning (MARL) provides an efficient way for simultaneously learning policies for multiple agents interacting with each other. However, in scenarios requiring complex interactions, existing algorithms can suffer from an inability to accurately anticipate the influence of self-actions on other agents. Incorporating an ability to reason about other agents' potential responses can allow an agent to formulate more effective strategies. This paper adopts a recursive reasoning model in a centralized-training-decentralized-execution framework to help learning agents better cooperate with or compete against others. The proposed algorithm, referred to as the Recursive Reasoning Graph (R2G), shows state-of-the-art performance on multiple multi-agent particle and robotics games.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 156f2093-cc16-4abf-bf32-5e774c883bc4Builds on1
Related papers
- Reinforcement Learning under a Multi-agent Predictive State Representation Model: Method and TheoryZhi Zhang, Zhuoran Yang, Han Liu, Pratap Tokekar et al.ICLR 2022 · 10 citations
- R2-B2: Recursive Reasoning-Based Bayesian Optimization for No-Regret Learning in GamesZhongxiang Dai, Yizhou Chen, Bryan Kian Hsiang Low, Patrick Jaillet et al.ICML 2020 · 28 citations
- Multi-Agent Actor-Critic with Hierarchical Graph Attention NetworkHeechang Ryu, Hayong Shin, Jinkyoo ParkAAAI 2020 · 143 citations
- Mimicking To Dominate: Imitation Learning Strategies for Success in Multiagent GamesThe Viet Bui, Tien Mai, Thanh Hong NguyenNeurIPS 2024 · 5 citations
- Policy Gradient With Serial Markov Chain ReasoningEdoardo Cetin, Oya ÇeliktutanNeurIPS 2022 · 4 citations
