Finding and Certifying (Near-)Optimal Strategies in Black-Box Extensive-Form Games
Brian Hu Zhang, Tuomas Sandholm
摘要
Often-for example in war games, strategy video games, and financial simulations-the game is given to us only as a black-box simulator in which we can play it. In these settings, since the game may have unknown nature action distributions (from which we can only obtain samples) and/or be too large to expand fully, it can be difficult to compute strategies with guarantees on exploitability. Recent work (Zhang and Sandholm 2020) resulted in a notion of certificate for extensive-form games that allows exploitability guarantees while not expanding the full game tree. However, that work assumed that the black box could sample or expand arbitrary nodes of the game tree at any time, and that a series of exact game solves (via, for example, linear programming) can be conducted to compute the certificate. Each of those two assumptions severely restricts the practical applicability of that method. In this work, we relax both of the assumptions. We show that high-probability certificates can be obtained with a black box that can do nothing more than play through games, using only a regret minimizer as a subroutine. As a bonus, we obtain an equilibrium-finding algorithm with Õ(1/ √ T ) convergence rate in the extensive-form game setting that does not rely on a sampling strategy with lower-bounded reach probabilities (which MCCFR assumes). We demonstrate experimentally that, in the black-box setting, our methods are able to provide nontrivial exploitability guarantees while expanding only a small fraction of the game tree. Introduction Computational equilibrium finding has led to many recent breakthroughs in AI in games such as poker (Bowling et al. 2015; Brown and Sandholm 2017; Moravčík et al. 2017; Brown and Sandholm 2019b) where the game is fully known. However, in many applications, the game is not fully known; instead, it is given only via a simulator that permits an algorithm to play through the game repeatedly (e.g., Wellman 2006; Lanctot et al. 2017; Tuyls et al. 2018; Areyan Viqueira, Cousins, and Greenwald 2020) . The algorithm may never know the game exactly. While deep reinforcement learning has yielded strong practical results in this
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Efficient Phi-Regret Minimization in Extensive-Form Games via Online Mirror DescentYu Bai, Chi Jin, Song Mei, Ziang Song 等NeurIPS 2022 · 被引用 24 次
- Learning in two-player zero-sum partially observable Markov games with perfect recallTadashi Kozuno, Pierre Ménard, Rémi Munos, Michal ValkoNeurIPS 2021 · 被引用 23 次
- Subgame solving without common knowledgeBrian Hu Zhang, Tuomas SandholmNeurIPS 2021 · 被引用 21 次
- Adapting to game trees in zero-sum imperfect information gamesCôme Fiegel, Pierre Ménard, Tadashi Kozuno, Rémi Munos 等ICML 2023 · 被引用 13 次
- Sample-Efficient Learning of Correlated Equilibria in Extensive-Form GamesZiang Song, Song Mei, Yu BaiNeurIPS 2022 · 被引用 11 次
它引用的顶会 Paper4
- Stochastic Regret Minimization in Extensive-Form GamesGabriele Farina, Christian Kroer, Tuomas SandholmICML 2020 · 被引用 32 次
- Bandit Linear Optimization for Sequential Decision Making and Extensive-Form GamesGabriele Farina, Robin Schmucker, Tuomas SandholmAAAI 2021 · 被引用 25 次
- Model-Free Online Learning in Unknown Sequential Decision Making Problems and GamesGabriele Farina, Tuomas SandholmAAAI 2021 · 被引用 24 次
- Small Nash Equilibrium Certificates in Very Large GamesBrian Hu Zhang, Tuomas SandholmNeurIPS 2020 · 被引用 6 次
相关 Paper
- Sparsified Linear Programming for Zero-Sum Equilibrium FindingBrian Hu Zhang, Tuomas SandholmICML 2020 · 被引用 11 次
- ESCHER: Eschewing Importance Sampling in Games by Computing a History Value Function to Estimate RegretStephen Marcus McAleer, Gabriele Farina, Marc Lanctot, Tuomas SandholmICLR 2023 · 被引用 1 次
- Global Policy-Space Response Oracles for Two-Player Zero-Sum GamesJunyu Zhang, Feihong Yang, Jian Wang, Chao Wang 等ICML 2026
- Polynomial-Time Optimal Equilibria with a Mediator in Extensive-Form GamesBrian Hu Zhang, Tuomas SandholmNeurIPS 2022 · 被引用 15 次
- Computing Optimal Equilibria and Mechanisms via Learning in Zero-Sum Extensive-Form GamesBrian Hu Zhang, Gabriele Farina, Ioannis Anagnostides, Federico Cacciamani 等NeurIPS 2023 · 被引用 17 次
