Global Policy-Space Response Oracles for Two-Player Zero-Sum Games
Junyu Zhang, Feihong Yang, Jian Wang, Chao Wang, Xudong Zhang
摘要
The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy set using deep reinforcement learning (DRL). A central challenge is to construct, under limited computational budgets, a small strategy population whose induced game well approximates the full game. Existing PSRO variants typically expand the population using best responses to meta-strategies computed from restricted-game payoffs, which can lead to inefficient expansions that provide limited global improvement. We propose to guide population expansion by directly evaluating the post-expansion population quality. Specifically, we adopt Population Exploitability (PE) to measure how well a restricted strategy set represents the full game, and introduce a two-phase exploration--selection framework that explicitly minimizes PE during expansion. We instantiate this framework as Global PSRO, a practical DRL-based algorithm that efficiently generates candidate responses and estimates PE via parameter-sharing conditional neural networks. Experiments across multiple two-player zero-sum games show that Global PSRO achieves lower exploitability and approximates Nash equilibria with significantly fewer policy iterations than prior PSRO methods. The experimental code is available at https://github.com/Zhangjy1997/GlobalPSRO.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- A Generalized Training Approach for Multiagent LearningPaul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls 等ICLR 2020 · 被引用 110 次
- Modelling Behavioural Diversity for Learning in Open-Ended GamesNicolas Perez Nieves, Yaodong Yang, Oliver Slumbers, David Henry Mguni 等ICML 2021 · 被引用 80 次
- Towards Unifying Behavioral and Response Diversity for Open-ended Learning in Zero-sum GamesXiangyu Liu, Hangtian Jia, Ying Wen, Yujing Hu 等NeurIPS 2021 · 被引用 67 次
- A Unified Diversity Measure for Multiagent Reinforcement LearningZongkai Liu, Chao Yu, Yaodong Yang, Peng Sun 等NeurIPS 2022 · 被引用 17 次
- Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum GamesSiqi Liu, Marc Lanctot, Luke Marris, Nicolas HeessICML 2022 · 被引用 12 次
相关 Paper
- XDO: A Double Oracle Algorithm for Extensive-Form GamesStephen McAleer, John B. Lanier, Kevin A. Wang, Pierre Baldi 等NeurIPS 2021 · 被引用 66 次
- Pipeline PSRO: A Scalable Approach for Finding Approximate Nash Equilibria in Large GamesStephen McAleer, John B. Lanier, Roy Fox, Pierre BaldiNeurIPS 2020 · 被引用 98 次
- Explicit Exploration for High-Welfare Equilibria in Game-Theoretic Multiagent Reinforcement LearningAustin A. Nguyen, Anri Gu, Michael P. WellmanICML 2025
- Toward Optimal Policy Population Growth in Two-Player Zero-Sum GamesStephen Marcus McAleer, JB Lanier, Kevin A. Wang, Pierre Baldi 等ICLR 2024 · 被引用 3 次
- Policy Space Diversity for Non-Transitive GamesJian Yao, Weiming Liu, Haobo Fu, Yaodong Yang 等NeurIPS 2023 · 被引用 28 次
