Neural Auto-Curricula in Two-Player Zero-Sum Games
Xidong Feng, Oliver Slumbers, Ziyu Wan, Bo Liu, Stephen McAleer, Ying Wen, Jun Wang, Yaodong Yang
摘要
When solving two-player zero-sum games, multi-agent reinforcement learning (MARL) algorithms often create populations of agents where, at each iteration, a new agent is discovered as the best response to a mixture over the opponent population. Within such a process, the update rules of "who to compete with" (i.e., the opponent mixture) and "how to beat them" (i.e., finding best responses) are underpinned by manually developed game theoretical principles such as fictitious play and Double Oracle. In this paper1, we introduce a novel framework-Neural Auto-Curricula (NAC)-that leverages meta-gradient descent to automate the discovery of the learning update rule without explicit human design. Specifically, we parameterise the opponent selection module by neural networks and the best-response module by optimisation subroutines, and update their parameters solely via interaction with the game engine, where both players aim to minimise their exploitability. Surprisingly, even without human design, the discovered MARL algorithms achieve competitive or even better performance with the state-of-the-art population-based game solvers (e.g., PSRO) on Games of Skill, differentiable Lotto, non-transitive Mixture Games, Iterated Matching Pennies, and Kuhn Poker. Additionally, we show that NAC is able to generalise from small games to large games, for example training on Kuhn Poker and outperforming PSRO on Leduc Poker. Our work inspires a promising future direction to discover general MARL algorithms solely from data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Discovered Policy OptimisationChris Lu, Jakub Grudzien Kuba, Alistair Letcher, Luke Metz 等NeurIPS 2022 · 被引用 134 次
- Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement LearningRunze Liu, Fengshuo Bai, Yali Du, Yaodong YangNeurIPS 2022 · 被引用 72 次
- Policy Space Diversity for Non-Transitive GamesJian Yao, Weiming Liu, Haobo Fu, Yaodong Yang 等NeurIPS 2023 · 被引用 28 次
- A Game-Theoretic Framework for Managing Risk in Multi-Agent SystemsOliver Slumbers, David Henry Mguni, Stefano B. Blumberg, Stephen Marcus McAleer 等ICML 2023 · 被引用 25 次
- Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium SolversLuke Marris, Ian Gemp, Thomas Anthony, Andrea Tacchetti 等NeurIPS 2022 · 被引用 22 次
相关 Paper
- Discovering Reinforcement Learning AlgorithmsJunhyuk Oh, Matteo Hessel, Wojciech M. Czarnecki, Zhongwen Xu 等NeurIPS 2020 · 被引用 154 次
- A Generalized Training Approach for Multiagent LearningPaul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls 等ICLR 2020 · 被引用 110 次
- Game-Theoretic Co-Evolution for LLM-Based Heuristic DiscoveryXinyi Ke, Kai Li, Junliang Xing, Yifan Zhang 等ICML 2026 · 被引用 1 次
- Discovering General Reinforcement Learning Algorithms with Adversarial Environment DesignMatthew Thomas Jackson, Minqi Jiang, Jack Parker-Holder, Risto Vuorio 等NeurIPS 2023 · 被引用 23 次
- Toward Optimal Policy Population Growth in Two-Player Zero-Sum GamesStephen Marcus McAleer, JB Lanier, Kevin A. Wang, Pierre Baldi 等ICLR 2024 · 被引用 3 次
