Scalable Online Planning via Reinforcement Learning Fine-Tuning
Arnaud Fickinger, Hengyuan Hu, Brandon Amos, Stuart J. Russell, Noam Brown
Abstract
Lookahead search has been a critical component of recent AI successes, such as in the games of chess, go, and poker. However, the search methods used in these games, and in many other settings, are tabular. Tabular search methods do not scale well with the size of the search space, and this problem is exacerbated by stochasticity and partial observability. In this work we replace tabular search with online model-based fine-tuning of a policy neural network via reinforcement learning, and show that this approach outperforms state-of-the-art search algorithms in benchmark settings. In particular, we use our search algorithm to achieve a new state-of-the-art result in self-play Hanabi, and show the generality of our algorithm by also showing that it outperforms tabular search in the Atari game Ms. Pacman.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b15e4006-28a6-4209-be06-94051c5ac03eCited by top-tier papers9
- Modeling Strong and Human-Like Gameplay with KL-Regularized SearchAthul Paul Jacob, David J. Wu, Gabriele Farina, Adam Lerer et al.ICML 2022 · 69 citations
- Thinker: Learning to Plan and ActStephen Chung, Ivan Anokhin, David KruegerNeurIPS 2023 · 18 citations
- A Fine-Tuning Approach to Belief State ModelingSamuel Sokota, Hengyuan Hu, David J. Wu, J. Zico Kolter et al.ICLR 2022 · 13 citations
- Abstracting Imperfect Information Away from Two-Player Zero-Sum GamesSamuel Sokota, Ryan D'Orazio, Chun Kai Ling, David J. Wu et al.ICML 2023 · 8 citations
- Test-time Offline Reinforcement Learning on Goal-related ExperienceMarco Bagatella, Mert Albaba, Jonas Hübotter, Georg Martius et al.ICML 2026 · 7 citations
Builds on11
- "Other-Play" for Zero-Shot CoordinationHengyuan Hu, Adam Lerer, Alex Peysakhovich, Jakob N. FoersterICML 2020 · 271 citations
- Combining Deep Reinforcement Learning and Search for Imperfect-Information GamesNoam Brown, Anton Bakhtin, Adam Lerer, Qucheng GongNeurIPS 2020 · 205 citations
- Exploring Model-based Planning with Policy NetworksTingwu Wang, Jimmy BaICLR 2020 · 164 citations
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain et al.NeurIPS 2021 · 149 citations
- Learning and Planning in Complex Action SpacesThomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain et al.ICML 2021 · 99 citations
Related papers
- The Update-Equivalence Framework for Decision-Time PlanningSamuel Sokota, Gabriele Farina, David J. Wu, Hengyuan Hu et al.ICLR 2024 · 5 citations
- Improving Policies via Search in Cooperative Partially Observable GamesAdam Lerer, Hengyuan Hu, Jakob N. Foerster, Noam BrownAAAI 2020 · 87 citations
- Equivariant Networks for Zero-Shot CoordinationDarius Muglich, Christian Schröder de Witt, Elise van der Pol, Shimon Whiteson et al.NeurIPS 2022 · 24 citations
- The Hidden Rules of Hanabi: How Humans Outperform AI AgentsMatthew Sidji, Wally Smith, Melissa J. RogersonCHI 2023 · 9 citations
- Muesli: Combining Improvements in Policy OptimizationMatteo Hessel, Ivo Danihelka, Fabio Viola, Arthur Guez et al.ICML 2021 · 69 citations
