Simplex Neural Population Learning: Any-Mixture Bayes-Optimality in Symmetric Zero-sum Games
Siqi Liu, Marc Lanctot, Luke Marris, Nicolas Heess
摘要
Learning to play optimally against any mixture over a diverse set of strategies is of important practical interests in competitive games. In this paper, we propose simplex-NeuPL that satisfies two desiderata simultaneously: i) learning a population of strategically diverse basis policies, represented by a single conditional network; ii) using the same network, learn best-responses to any mixture over the simplex of basis policies. We show that the resulting conditional policies incorporate prior information about their opponents effectively, enabling near optimal returns against arbitrary mixture policies in a game with tractable best-responses. We verify that such policies behave Bayes-optimally under uncertainty and offer insights in using this flexibility at test time. Finally, we offer evidence that learning best-responses to any mixture policies is an effective auxiliary task for strategic exploration, which, by itself, can lead to more performant populations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Turbocharging Solution Concepts: Solving NEs, CEs and CCEs with Neural Equilibrium SolversLuke Marris, Ian Gemp, Thomas Anthony, Andrea Tacchetti 等NeurIPS 2022 · 被引用 22 次
- Computing Optimal Equilibria and Mechanisms via Learning in Zero-Sum Extensive-Form GamesBrian Hu Zhang, Gabriele Farina, Ioannis Anagnostides, Federico Cacciamani 等NeurIPS 2023 · 被引用 17 次
- Multi-Agent Meta-Reinforcement Learning: Sharper Convergence Rates with Task SimilarityWeichao Mao, Haoran Qiu, Chen Wang, Hubertus Franke 等NeurIPS 2023 · 被引用 17 次
- Are Equivariant Equilibrium Approximators Beneficial?Zhijian Duan, Yunxuan Ma, Xiaotie DengICML 2023 · 被引用 4 次
- NfgTransformer: Equivariant Representation Learning for Normal-form GamesSiqi Liu, Luke Marris, Georgios Piliouras, Ian Gemp 等ICLR 2024 · 被引用 2 次
它引用的顶会 Paper6
- VariBAD: A Very Good Method for Bayes-Adaptive Deep RL via Meta-LearningLuisa M. Zintgraf, Kyriacos Shiarlis, Maximilian Igl, Sebastian Schulze 等ICLR 2020 · 被引用 315 次
- Meta-trained agents implement Bayes-optimal agentsVladimir Mikulik, Grégoire Delétang, Tom McGrath, Tim Genewein 等NeurIPS 2020 · 被引用 56 次
- OPtions as REsponses: Grounding behavioural hierarchies in multi-agent reinforcement learningAlexander Vezhnevets, Yuhuai Wu, Maria K. Eckstein, Rémi Leblond 等ICML 2020 · 被引用 44 次
- Neural Auto-Curricula in Two-Player Zero-Sum GamesXidong Feng, Oliver Slumbers, Ziyu Wan, Bo Liu 等NeurIPS 2021 · 被引用 40 次
- Iterative Empirical Game Solving via Single Policy Best ResponseMax Olan Smith, Thomas Anthony, Michael P. WellmanICLR 2021 · 被引用 23 次
相关 Paper
- NeuPL: Neural Population LearningSiqi Liu, Luke Marris, Daniel Hennes, Josh Merel 等ICLR 2022 · 被引用 19 次
- Generalization in Mean Field Games by Learning Master PoliciesSarah Perrin, Mathieu Laurière, Julien Pérolat, Romuald Élie 等AAAI 2022 · 被引用 47 次
- Evolution Strategies for Approximate Solution of Bayesian GamesZun Li, Michael P. WellmanAAAI 2021 · 被引用 19 次
- In-Context Learning Strategies Emerge RationallyDaniel Wurgaft, Ekdeep Singh Lubana, Core Francisco Park, Hidenori Tanaka 等NeurIPS 2025 · 被引用 19 次
- Generating Diverse Cooperative Agents by Learning Incompatible PoliciesRujikorn Charakorn, Poramate Manoonpong, Nat DilokthanakulICLR 2023
