A Policy Gradient Algorithm for Learning to Learn in Multiagent Reinforcement Learning
Dong-Ki Kim, Miao Liu, Matthew Riemer, Chuangchuang Sun, Marwa Abdulhai, Golnaz Habibi, Sebastian Lopez-Cot, Gerald Tesauro, Jonathan P. How
摘要
A fundamental challenge in multiagent reinforcement learning is to learn beneficial behaviors in a shared environment with other simultaneously learning agents. In particular, each agent perceives the environment as effectively nonstationary due to the changing policies of other agents. Moreover, each agent is itself constantly learning, leading to natural non-stationarity in the distribution of experiences encountered. In this paper, we propose a novel meta-multiagent policy gradient theorem that directly accounts for the non-stationary policy dynamics inherent to multiagent learning settings. This is achieved by modeling our gradient updates to consider both an agent's own non-stationary policy dynamics and the non-stationary policy dynamics of other agents in the environment. We show that our theoretically grounded approach provides a general solution to the multiagent learning problem, which inherently comprises all key aspects of previous state of the art approaches on this topic. We test our method on a diverse suite of multiagent benchmarks and demonstrate a more efficient ability to adapt to new agents as they learn than baseline methods across the full spectrum of mixed incentive, competitive, and cooperative domains. Our contributions. With this insight, we make the following primary contributions in this paper: 1) New theorem: We derive a new meta-multiagent policy gradient theorem (Meta-MAPG) that, for the first time,
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Model-Based Opponent ModelingXiaopeng Yu, Jiechuan Jiang, Wanpeng Zhang, Haobin Jiang 等NeurIPS 2022 · 被引用 56 次
- Model-Free Opponent ShapingChristopher Lu, Timon Willi, Christian A. Schröder de Witt, Jakob N. FoersterICML 2022 · 被引用 53 次
- Influencing Long-Term Behavior in Multiagent Reinforcement LearningDong-Ki Kim, Matthew Riemer, Miao Liu, Jakob N. Foerster 等NeurIPS 2022 · 被引用 29 次
- Continual Learning In Environments With Polynomial Mixing TimesMatthew Riemer, Sharath Chandra Raparthy, Ignacio Cases, Gopeshh Subbaraj 等NeurIPS 2022 · 被引用 18 次
- A Theoretical Understanding of Gradient Bias in Meta-Reinforcement LearningBo Liu, Xidong Feng, Jie Ren, Luo Mai 等NeurIPS 2022 · 被引用 16 次
它引用的顶会 Paper1
相关 Paper
- Model-based Adversarial Meta-Reinforcement LearningZichuan Lin, Garrett Thomas, Guangwen Yang, Tengyu MaNeurIPS 2020 · 被引用 58 次
- Can Learned Optimization Make Reinforcement Learning Less Difficult?Alexander David Goldie, Chris Lu, Matthew Thomas Jackson, Shimon Whiteson 等NeurIPS 2024 · 被引用 18 次
- Distributional Meta-Gradient Reinforcement LearningHaiyan Yin, Shuicheng Yan, Zhongwen XuICLR 2023
- On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement LearningAlireza Fallah, Kristian Georgiev, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2021 · 被引用 31 次
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh 等NeurIPS 2020 · 被引用 90 次
