Policy Space Diversity for Non-Transitive Games
Jian Yao, Weiming Liu, Haobo Fu, Yaodong Yang, Stephen McAleer, Qiang Fu, Wei Yang
Abstract
Policy-Space Response Oracles (PSRO) is an influential algorithm framework for approximating a Nash Equilibrium (NE) in multi-agent non-transitive games. Many previous studies have been trying to promote policy diversity in PSRO. A major weakness in existing diversity metrics is that a more diverse (according to their diversity metrics) population does not necessarily mean (as we proved in the paper) a better approximation to a NE. To alleviate this problem, we propose a new diversity metric, the improvement of which guarantees a better approximation to a NE. Meanwhile, we develop a practical and well-justified method to optimize our diversity metric using only state-action samples. By incorporating our diversity regularization into the best response solving in PSRO, we obtain a new PSRO variant, Policy Space Diversity PSRO (PSD-PSRO). We present the convergence property of PSD-PSRO. Empirically, extensive experiments on various games demonstrate that PSD-PSRO is more effective in producing significantly less exploitable policies than state-of-the-art PSRO variants. The experiment code is available at https://github.com/nigelyaoj/policy-space-diversity-psro .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Diversity-Aware Policy Optimization for Large Language Model ReasoningJian Yao, Ran Cheng, Xingyu Wu, Jibin Wu et al.NeurIPS 2025 · 44 citations
- Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-PlayQinsi Wang, Bo Liu, Tianyi Zhou, Jing Shi et al.ICLR 2026 · 24 citations
- Team-PSRO for Learning Approximate TMECor in Large Team Games via Cooperative Reinforcement LearningStephen McAleer, Gabriele Farina, Gaoyue Zhou, Mingzhi Wang et al.NeurIPS 2023 · 18 citations
- Reevaluating Policy Gradient Methods for Imperfect-Information GamesMax Rudolph, Nathan Lichtlé, Sobhan Mohammadpour, Alexandre M Bayen et al.ICLR 2026 · 17 citations
- HM3: Hierarchical Multi-Objective Model Merging for Pretrained ModelsYu Zhou, Xingyu Wu, Jibin Wu, Liang Feng et al.NeurIPS 2025 · 14 citations
Builds on19
- Effective Diversity in Population Based Reinforcement LearningJack Parker-Holder, Aldo Pacchiano, Krzysztof Marcin Choromanski, Stephen J. RobertsNeurIPS 2020 · 195 citations
- Trajectory Diversity for Zero-Shot CoordinationAndrei Lupu, Brandon Cui, Hengyuan Hu, Jakob N. FoersterICML 2021 · 157 citations
- Real World Games Look Like Spinning TopsWojciech M. Czarnecki, Gauthier Gidel, Brendan D. Tracey, Karl Tuyls et al.NeurIPS 2020 · 123 citations
- One Solution is Not All You Need: Few-Shot Extrapolation via Structured MaxEnt RLSaurabh Kumar, Aviral Kumar, Sergey Levine, Chelsea FinnNeurIPS 2020 · 109 citations
- Pipeline PSRO: A Scalable Approach for Finding Approximate Nash Equilibria in Large GamesStephen McAleer, John B. Lanier, Roy Fox, Pierre BaldiNeurIPS 2020 · 98 citations
Related papers
- A Unified Diversity Measure for Multiagent Reinforcement LearningZongkai Liu, Chao Yu, Yaodong Yang, Peng Sun et al.NeurIPS 2022 · 17 citations
- A-PSRO: A Unified Strategy Learning Method with Advantage Metric for Normal-form GamesYudong Hu, Haoran Li, Congying Han, Tiande Guo et al.ICML 2025
- Global Policy-Space Response Oracles for Two-Player Zero-Sum GamesJunyu Zhang, Feihong Yang, Jian Wang, Chao Wang et al.ICML 2026
- Explicit Exploration for High-Welfare Equilibria in Game-Theoretic Multiagent Reinforcement LearningAustin A. Nguyen, Anri Gu, Michael P. WellmanICML 2025
- Iterative Empirical Game Solving via Single Policy Best ResponseMax Olan Smith, Thomas Anthony, Michael P. WellmanICLR 2021 · 23 citations
