Twice Sequential Monte Carlo for Tree Search
Yaniv Oren, Joery de Vries, Pascal Van der Vaart, Matthijs T. J. Spaan, Wendelin Boehmer
Abstract
Model-based reinforcement learning (RL) methods that leverage search are responsible for many milestone breakthroughs in RL. Sequential Monte Carlo (SMC) recently emerged as an alternative to the Monte Carlo Tree Search (MCTS) algorithm which drove these breakthroughs. SMC is easier to parallelize and more suitable to GPU acceleration. However, it also suffers from large variance and path degeneracy which prevent it from scaling well with increased search depth, i.e., increased sequential compute. To address these problems, we introduce Twice Sequential Monte Carlo Tree Search (TSMCTS). Across discrete and continuous environments TSMCTS outperforms the SMC baseline as well as a popular modern version of MCTS as a policy improvement operator, scales favorably with sequential compute, reduces estimator variance and mitigates the effects of path degeneracy while retaining the properties that make SMC natural to parallelize.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 294ed8fa-e44e-49f8-86f9-f967fae43b52Builds on12
- Online and Offline Reinforcement Learning by Planning with a Learned ModelJulian Schrittwieser, Thomas Hubert, Amol Mandhane, Mohammadamin Barekatain et al.NeurIPS 2021 · 149 citations
- Learning and Planning in Complex Action SpacesThomas Hubert, Julian Schrittwieser, Ioannis Antonoglou, Mohammadamin Barekatain et al.ICML 2021 · 99 citations
- Policy improvement by planning with GumbelIvo Danihelka, Arthur Guez, Julian Schrittwieser, David SilverICLR 2022 · 84 citations
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert et al.ICML 2020 · 84 citations
- Jumanji: a Diverse Suite of Scalable Reinforcement Learning Environments in JAXClément Bonnet, Daniel Luo, Donal Byrne, Shikha Surana et al.ICLR 2024 · 52 citations
Related papers
- Trust-Region Twisted Policy ImprovementJoery A. de Vries, Jinke He, Yaniv Oren, Matthijs T. J. SpaanICML 2025
- SPO: Sequential Monte Carlo Policy OptimisationMatthew Macfarlane, Edan Toledo, Donal Byrne, Paul Duckworth et al.NeurIPS 2024 · 8 citations
- Watch the Unobserved: A Simple Approach to Parallelizing Monte Carlo Tree SearchAnji Liu, Jianshu Chen, Mingze Yu, Yu Zhai et al.ICLR 2020 · 39 citations
- Efficient Learning of Deep State Space Models via Importance SmoothingJohn-Joseph Brady, Nikolas Nüsken, Yunpeng LiICML 2026
- Practical Massively Parallel Monte-Carlo Tree Search Applied to Molecular DesignXiufeng Yang, Tanuj Kr Aasawat, Kazuki YoshizoeICLR 2021 · 25 citations
