Guarantees for Self-Play in Multiplayer Games via Polymatrix Decomposability
Revan MacQueen, James R. Wright
摘要
Self-play is a technique for machine learning in multi-agent systems where a learning algorithm learns by interacting with copies of itself. Self-play is useful for generating large quantities of data for learning, but has the drawback that the agents the learner will face post-training may have dramatically different behavior than the learner came to expect by interacting with itself. For the special case of two-player constant-sum games, self-play that reaches Nash equilibrium is guaranteed to produce strategies that perform well against any post-training opponent; however, no such guarantee exists for multiplayer games. We show that in games that approximately decompose into a set of two-player constant-sum games (called constant-sum polymatrix games) where global -Nash equilibria are boundedly far from Nash equilibria in each subgame (called subgame stability), any no-external-regret algorithm that learns by self-play will produce a strategy with bounded vulnerability. For the first time, our results identify a structural property of multiplayer games that enable performance guarantees for the strategies produced by a broad class of self-play algorithms. We demonstrate our findings through experiments on Leduc poker.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Why Playing Against Diverse and Challenging Opponents Speeds Up Coevolution: A Theoretical Analysis on Combinatorial GamesAlistair Benford, Per Kristian LehreNeurIPS 2025 · 被引用 2 次
- Exploiting Structure in Offline Multi-Agent RL: The Benefits of Low Interaction RankWenhao Zhan, Scott Fujimoto, Zheqing Zhu, Jason D. Lee 等ICLR 2025
它引用的顶会 Paper9
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 被引用 271 次
- A Sharp Analysis of Model-based Reinforcement Learning with Self-PlayQinghua Liu, Tiancheng Yu, Yu Bai, Chi JinICML 2021 · 被引用 137 次
- No-Press Diplomacy from ScratchAnton Bakhtin, David J. Wu, Adam Lerer, Noam BrownNeurIPS 2021 · 被引用 51 次
- No-Regret Learning Dynamics for Extensive-Form Correlated EquilibriumAndrea Celli, Alberto Marchesi, Gabriele Farina, Nicola GattiNeurIPS 2020 · 被引用 48 次
- Multi-Agent Training beyond Zero-Sum with Correlated Equilibrium Meta-SolversLuke Marris, Paul Muller, Marc Lanctot, Karl Tuyls 等ICML 2021 · 被引用 42 次
相关 Paper
- Securing Equal Share: A Principled Approach for Learning Multiplayer Symmetric GamesJiawei Ge, Yuanhao Wang, Wenzhe Li, Chi JinICML 2025
- Convergence of No-Swap-Regret Dynamics in Self-PlayRenato Paes Leme, Georgios Piliouras, Jon SchneiderNeurIPS 2024 · 被引用 3 次
- For Learning in Symmetric Teams, Local Optima are Global Nash EquilibriaScott Emmons, Caspar Oesterheld, Andrew Critch, Vincent Conitzer 等ICML 2022 · 被引用 12 次
- Data Poisoning to Fake a Nash Equilibria for Markov GamesYoung Wu, Jeremy McMahan, Xiaojin Zhu, Qiaomin XieAAAI 2024 · 被引用 6 次
- Meta-Learning in GamesKeegan Harris, Ioannis Anagnostides, Gabriele Farina, Mikhail Khodak 等ICLR 2023 · 被引用 196 次
