Toward Optimal Policy Population Growth in Two-Player Zero-Sum Games
Stephen Marcus McAleer, JB Lanier, Kevin A. Wang, Pierre Baldi, Tuomas Sandholm, Roy Fox
Abstract
In competitive two-agent environments, deep reinforcement learning (RL) methods based on the Double Oracle (DO) algorithm, such as Policy Space Response Oracles (PSRO) and Anytime PSRO (APSRO), iteratively add RL best response policies to a population. Eventually, an optimal mixture of these population policies will approximate a Nash equilibrium. However, these methods might need to add all deterministic policies before converging. In this work, we introduce Self-Play PSRO (SP-PSRO), a method that adds an approximately optimal stochastic policy to the population in each iteration. Instead of adding only deterministic best responses to the opponent's least exploitable population mixture, SP-PSRO also learns an approximately optimal stochastic policy and adds it to the population as well. As a result, SP-PSRO empirically tends to converge much faster than APSRO and in many games converges in just a few iterations. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f6438889-c511-49e3-9194-b6ebea9a51f5Cited by top-tier papers2
- Reevaluating Policy Gradient Methods for Imperfect-Information GamesMax Rudolph, Nathan Lichtlé, Sobhan Mohammadpour, Alexandre M Bayen et al.ICLR 2026 · 17 citations
- Battling against Tough Resister: Strategy Planning with Adversarial Game for Non-collaborative DialoguesHaiyang Wang, Zhiliang Tian, Yuchen Pan, Xin Song et al.ACL 2025 · 3 citations
Builds on13
- Independent Policy Gradient Methods for Competitive Reinforcement LearningConstantinos Daskalakis, Dylan J. Foster, Noah GolowichNeurIPS 2020 · 200 citations
- Global Convergence of Multi-Agent Policy Gradient in Markov Potential GamesStefanos Leonardos, Will Overman, Ioannis Panageas, Georgios PiliourasICLR 2022 · 158 citations
- A Generalized Training Approach for Multiagent LearningPaul Muller, Shayegan Omidshafiei, Mark Rowland, Karl Tuyls et al.ICLR 2020 · 110 citations
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via RegularizationJulien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei et al.ICML 2021 · 102 citations
- Pipeline PSRO: A Scalable Approach for Finding Approximate Nash Equilibria in Large GamesStephen McAleer, John B. Lanier, Roy Fox, Pierre BaldiNeurIPS 2020 · 98 citations
Related papers
- Iterative Empirical Game Solving via Single Policy Best ResponseMax Olan Smith, Thomas Anthony, Michael P. WellmanICLR 2021 · 23 citations
- XDO: A Double Oracle Algorithm for Extensive-Form GamesStephen McAleer, John B. Lanier, Kevin A. Wang, Pierre Baldi et al.NeurIPS 2021 · 66 citations
- Global Policy-Space Response Oracles for Two-Player Zero-Sum GamesJunyu Zhang, Feihong Yang, Jian Wang, Chao Wang et al.ICML 2026
- Explicit Exploration for High-Welfare Equilibria in Game-Theoretic Multiagent Reinforcement LearningAustin A. Nguyen, Anri Gu, Michael P. WellmanICML 2025
- Policy Space Diversity for Non-Transitive GamesJian Yao, Weiming Liu, Haobo Fu, Yaodong Yang et al.NeurIPS 2023 · 28 citations
