The Update-Equivalence Framework for Decision-Time Planning
Samuel Sokota, Gabriele Farina, David J. Wu, Hengyuan Hu, Kevin A. Wang, J. Zico Kolter, Noam Brown
摘要
The process of revising (or constructing) a policy at execution time -- known as decision-time planning -- has been key to achieving superhuman performance in perfect-information games like chess and Go. A recent line of work has extended decision-time planning to imperfect-information games, leading to superhuman performance in poker. However, these methods involve solving subgames whose sizes grow quickly in the amount of non-public information, making them unhelpful when the amount of non-public information is large. Motivated by this issue, we introduce an alternative framework for decision-time planning that is not based on solving subgames, but rather on update equivalence. In this update-equivalence framework, decision-time planning algorithms replicate the updates of last-iterate algorithms, which need not rely on public information. This facilitates scalability to games with large amounts of non-public information. Using this framework, we derive a provably sound search algorithm for fully cooperative games based on mirror descent and a search algorithm for adversarial games based on magnetic mirror descent. We validate the performance of these algorithms in cooperative and adversarial domains, notably in Hanabi, the standard benchmark for search in fully cooperative imperfect-information games. Here, our mirror descent approach exceeds or matches the performance of public information-based search while using two orders of magnitude less search time. This is the first instance of a non-public-information-based algorithm outperforming public-information-based approaches in a domain they have historically dominated.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- General search techniques without common knowledge for imperfect-information games, and application to superhuman Fog of War chessBrian Zhang, Tuomas SandholmICLR 2026 · 被引用 13 次
- Opponent Modeling with In-context SearchYuheng Jing, Bingyun Liu, Kai Li, Yifan Zang 等NeurIPS 2024 · 被引用 8 次
- Towards Offline Opponent Modeling with In-context LearningYuheng Jing, Kai Li, Bingyun Liu, Yifan Zang 等ICLR 2024 · 被引用 7 次
它引用的顶会 Paper13
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 被引用 205 次
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via RegularizationJulien Pérolat, Rémi Munos, Jean-Baptiste Lespiau, Shayegan Omidshafiei 等ICML 2021 · 被引用 102 次
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 被引用 87 次
- Planning in Stochastic Environments with a Learned ModelIoannis Antonoglou, Julian Schrittwieser, Sherjil Ozair, Thomas K. Hubert 等ICLR 2022 · 被引用 79 次
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez 等ICLR 2021 · 被引用 77 次
相关 Paper
- The Hidden Rules of Hanabi: How Humans Outperform AI AgentsMatthew Sidji, Wally Smith, Melissa J. RogersonCHI 2023 · 被引用 9 次
- Scalable Online Planning via Reinforcement Learning Fine-TuningArnaud Fickinger, Hengyuan Hu, Brandon Amos, Stuart J. Russell 等NeurIPS 2021 · 被引用 26 次
- Equivariant Networks for Zero-Shot CoordinationDarius Muglich, Christian Schröder de Witt, Elise van der Pol, Shimon Whiteson 等NeurIPS 2022 · 被引用 24 次
- Reevaluating Policy Gradient Methods for Imperfect-Information GamesMax Rudolph, Nathan Lichtlé, Sobhan Mohammadpour, Alexandre M Bayen 等ICLR 2026 · 被引用 17 次
- Abstracting Imperfect Information Away from Two-Player Zero-Sum GamesSamuel Sokota, Ryan D'Orazio, Chun Kai Ling, David J. Wu 等ICML 2023 · 被引用 8 次
