Reevaluating Policy Gradient Methods for Imperfect-Information Games
Max Rudolph, Nathan Lichtlé, Sobhan Mohammadpour, Alexandre M Bayen, J Zico Kolter, Amy Zhang, Gabriele Farina, Eugene Vinitsky, Samuel Sokota
Abstract
In the past decade, motivated by the putative failure of naive self-play deep reinforcement learning (DRL) in adversarial imperfect-information games, researchers have developed numerous DRL algorithms based on fictitious play (FP), double oracle (DO), and counterfactual regret minimization (CFR). In light of recent results of the magnetic mirror descent algorithm, we hypothesize that simpler generic policy gradient methods like PPO are competitive with or superior to these FP-, DO-, and CFR-based DRL approaches. To facilitate the resolution of this hypothesis, we implement and release the first broadly accessible exact exploitability computations for five large games. Using these games, we conduct the largest-ever exploitability comparison of DRL algorithms for imperfect-information games. Over 7000 training runs, we find that FP-, DO-, and CFR-based approaches fail to outperform generic policy gradient methods. Recently, Sokota et al. (2023) demonstrated the promise of an alternative algorithm, a policy gradient (PG) method called magnetic mirror descent (MMD).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Policy Gradient Methods Converge Globally in Imperfect-Information Extensive-Form GamesFivos Kalogiannis, Gabriele FarinaNeurIPS 2025 · 2 citations
- Solving Football by Exploiting Equilibrium Structure of 2p0s Differential Games with One-Sided InformationMukesh Ghimire, Lei Zhang, Zhe Xu, Yi RenICLR 2026 · 2 citations
- Provably Convergent Actor-Critic in Risk-averse MARLYizhou Zhang, Eric MazumdarICML 2026
Builds on15
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 205 citations
- Phasic Policy GradientKarl Cobbe, Jacob Hilton, Oleg Klimov, John SchulmanICML 2021 · 191 citations
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via RegularizationJulien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei et al.ICML 2021 · 102 citations
- Pipeline PSRO: A Scalable Approach for Finding Approximate Nash Equilibria in Large GamesStephen McAleer, John B. Lanier, Roy Fox, Pierre BaldiNeurIPS 2020 · 98 citations
Related papers
- Global Policy-Space Response Oracles for Two-Player Zero-Sum GamesJunyu Zhang, Feihong Yang, Jian Wang, Chao Wang et al.ICML 2026
- Policy Mirror Descent with LookaheadKimon Protopapas, Anas BarakatNeurIPS 2024 · 7 citations
- A Unified Approach to Reinforcement Learning, Quantal Response Equilibria, and Two-Player Zero-Sum GamesSamuel Sokota, Ryan D'Orazio, J. Zico Kolter, Nicolas Loizou et al.ICLR 2023 · 6 citations
- An Efficient Deep Reinforcement Learning Algorithm for Solving Imperfect Information Extensive-Form GamesLinjian Meng, Zhenxing Ge, Pinzhuo Tian, Bo An et al.AAAI 2023 · 8 citations
- Constrained Exploitability Descent: An Offline Reinforcement Learning Method for Finding Mixed-Strategy Nash EquilibriumRunyu Lu, Yuanheng Zhu, Dongbin ZhaoICML 2025
