XDO: A Double Oracle Algorithm for Extensive-Form Games
Stephen McAleer, John B. Lanier, Kevin A. Wang, Pierre Baldi, Roy Fox
摘要
Policy Space Response Oracles (PSRO) is a reinforcement learning (RL) algorithm for two-player zero-sum games that has been empirically shown to find approximate Nash equilibria in large games. Although PSRO is guaranteed to converge to an approximate Nash equilibrium and can handle continuous actions, it may take an exponential number of iterations as the number of information states (infostates) grows. We propose Extensive-Form Double Oracle (XDO), an extensive-form double oracle algorithm for two-player zero-sum games that is guaranteed to converge to an approximate Nash equilibrium linearly in the number of infostates. Unlike PSRO, which mixes best responses at the root of the game, XDO mixes best responses at every infostate. We also introduce Neural XDO (NXDO), where the best response is learned through deep RL. In tabular experiments on Leduc poker, we find that XDO achieves an approximate Nash equilibrium in a number of iterations an order of magnitude smaller than PSRO. Experiments on a modified Leduc poker game and Oshi-Zumo show that tabular XDO achieves a lower exploitability than CFR with the same amount of computation. We also find that NXDO outperforms PSRO and NFSP on a sequential multidimensional continuous-action game. NXDO is the first deep RL method that can find an approximate Nash equilibrium in high-dimensional continuous-action sequential games. Experiment code is available at https://github.com/indylab/nxdo .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-SolversLuke Marris, Paul Muller, Marc Lanctot, Karl Tuyls 等ICML 2021 · 被引用 42 次
- A Game-Theoretic Framework for Managing Risk in Multi-Agent SystemsOliver Slumbers, David Henry Mguni, Stefano B. Blumberg, Stephen Marcus McAleer 等ICML 2023 · 被引用 25 次
- Team-PSRO for Learning Approximate TMECor in Large Team Games via Cooperative Reinforcement LearningStephen McAleer, Gabriele Farina, Gaoyue Zhou, Mingzhi Wang 等NeurIPS 2023 · 被引用 18 次
- Computing Optimal Equilibria and Mechanisms via Learning in Zero-Sum Extensive-Form GamesBrian Hu Zhang, Gabriele Farina, Ioannis Anagnostides, Federico Cacciamani 等NeurIPS 2023 · 被引用 17 次
- Reevaluating Policy Gradient Methods for Imperfect-Information GamesMax Rudolph, Nathan Lichtlé, Sobhan Mohammadpour, Alexandre M Bayen 等ICLR 2026 · 被引用 17 次
它引用的顶会 Paper3
- Pipeline PSRO: A Scalable Approach for Finding Approximate Nash Equilibria in Large GamesStephen McAleer, John B. Lanier, Roy Fox, Pierre BaldiNeurIPS 2020 · 被引用 98 次
- Evolutionary Reinforcement Learning for Sample-Efficient Multiagent CoordinationSomdeb Majumdar, Shauharda Khadka, Santiago Miret, Stephen McAleer 等ICML 2020 · 被引用 70 次
- Double Neural Counterfactual Regret MinimizationHui Li, Kailiang Hu, Shaohua Zhang, Yuan Qi 等ICLR 2020 · 被引用 54 次
相关 Paper
- Global Policy-Space Response Oracles for Two-Player Zero-Sum GamesJunyu Zhang, Feihong Yang, Jian Wang, Chao Wang 等ICML 2026
- Toward Optimal Policy Population Growth in Two-Player Zero-Sum GamesStephen Marcus McAleer, JB Lanier, Kevin A. Wang, Pierre Baldi 等ICLR 2024 · 被引用 3 次
- Explicit Exploration for High-Welfare Equilibria in Game-Theoretic Multiagent Reinforcement LearningAustin A. Nguyen, Anri Gu, Michael P. WellmanICML 2025
- Iterative Empirical Game Solving via Single Policy Best ResponseMax Olan Smith, Thomas Anthony, Michael P. WellmanICLR 2021 · 被引用 23 次
- A Generalized Training Approach for Multiagent LearningPaul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls 等ICLR 2020 · 被引用 110 次
