Epistemic Monte Carlo Tree Search
Yaniv Oren, Viliam Vadocz, Matthijs T. J. Spaan, Wendelin Boehmer
Abstract
The AlphaZero/MuZero (A/MZ) family of algorithms has achieved remarkable success across various challenging domains by integrating Monte Carlo Tree Search (MCTS) with learned models. Learned models introduce epistemic uncertainty, which is caused by learning from limited data and is useful for exploration in sparse reward environments. MCTS does not account for the propagation of this uncertainty however. To address this, we introduce Epistemic MCTS (EMCTS): a theoretically motivated approach to account for the epistemic uncertainty in search and harness the search for deep exploration. In the challenging sparse-reward task of writing code in the Assembly language subleq, AZ paired with our method achieves significantly higher sample efficiency over baseline AZ. Search with EMCTS solves variations of the commonly used hard-exploration benchmark Deep Sea -which baseline A/MZ are practically unable to solve -much faster than an otherwise equivalent method that does not use search for uncertainty estimation, demonstrating significant benefits from search for epistemic uncertainty estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 72e031c1-55f8-4272-967b-46670d2a5a7dCited by top-tier papers5
- PlanU: Large Language Model Reasoning through Planning under UncertaintyZiwei Deng, Mian Deng, Chenjing Liang, Zeming Gao et al.NeurIPS 2025 · 4 citations
- Compositional Imprecise Probability: A Solution from Graded Monads and Markov CategoriesJack Liell-Cock, Sam StatonPOPL 2025 · 3 citations
- Twice Sequential Monte Carlo for Tree SearchYaniv Oren, Joery de Vries, Pascal Van der Vaart, Matthijs T. J. Spaan et al.ICML 2026 · 2 citations
- Trust-Region Twisted Policy ImprovementJoery A. de Vries, Jinke He, Yaniv Oren, Matthijs T. J. SpaanICML 2025
- Uncertainty-Aware Test-Time Search for Optimization Problem SolvingLinlin Yu, Xujiang Zhao, Dong Li, Yanchi Liu et al.ACL 2026
Builds on14
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
- Mastering Atari Games with Limited DataWeirui Ye, Shaohuai Liu, Thanard Kurutach, Pieter Abbeel et al.NeurIPS 2021 · 345 citations
- SUNRISE: A Simple Unified Framework for Ensemble Learning in Deep Reinforcement LearningKimin Lee, Michael Laskin, Aravind Srinivas, Pieter AbbeelICML 2021 · 239 citations
- Uncertainty Weighted Actor-Critic for Offline Reinforcement LearningYue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind et al.ICML 2021 · 223 citations
Related papers
- Monte-Carlo Tree Search as Regularized Policy OptimizationJean-Bastien Grill, Florent Altché, Yunhao Tang, Thomas Hubert et al.ICML 2020 · 84 citations
- Accelerating Monte Carlo Tree Search with Probability Tree State AbstractionYangqing Fu, Ming Sun, Buqing Nie, Yue GaoNeurIPS 2023 · 5 citations
- Goal-Directed Planning via Hindsight Experience ReplayLorenzo Moro, Amarildo Likmeta, Enrico Prati, Marcello RestelliICLR 2022 · 14 citations
- Efficient Exploration via Epistemic-Risk-Seeking Policy OptimizationBrendan O'DonoghueICML 2023 · 11 citations
- Learning to Stop: Dynamic Simulation Monte-Carlo Tree SearchLi-Cheng Lan, Ti-Rong Wu, I-Chen Wu, Cho-Jui HsiehAAAI 2021 · 7 citations
