Bayesian Optimized Monte Carlo Planning
John Mern, Anil Yildiz, Zachary Sunberg, Tapan Mukerji, Mykel J. Kochenderfer
Abstract
Online solvers for partially observable Markov decision processes have difficulty scaling to problems with large action spaces. Monte Carlo tree search with progressive widening attempts to improve scaling by sampling from the action space to construct a policy search tree. The performance of progressive widening search is dependent upon the action sampling policy, often requiring problem-specific samplers. In this work, we present a general method for efficient action sampling based on Bayesian optimization. The proposed method uses a Gaussian process to model a belief over the action-value function and selects the action that will maximize the expected improvement in the optimal action value. We implement the proposed approach in a new online tree search algorithm called Bayesian Optimized Monte Carlo Planning (BOMCP). Several experiments show that BOMCP is better able to scale to large action space POMDPs than existing state-of-the-art tree search solvers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5461c0d8-c7f4-4915-a821-c2a28748450dCited by top-tier papers6
- Risk-Averse Bayes-Adaptive Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2021 · 50 citations
- Information-guided Planning: An Online Approach for Partially Observable ProblemsMatheus Aparecido do Carmo Alves, Amokh Varma, Yehia Elkhatib, Leandro Soriano MarcolinoNeurIPS 2023 · 2 citations
- A Bayesian Approach to Online PlanningNir Greshler, David Ben-Eli, Carmel Rabinovitz, Gabi Guetta et al.ICML 2024 · 1 citation
- ReqsMiner: Automated Discovery of CDN Forwarding Request Inconsistencies and DoS Attacks with Grammar-based FuzzingLinkai Zheng, Xiang Li, Chuhan Wang, Run Guo et al.NDSS 2024
- Epistemic Monte Carlo Tree SearchYaniv Oren, Viliam Vadocz, Matthijs T. J. Spaan, Wendelin BoehmerICLR 2025
Builds on2
- Monte Carlo Tree Search in Continuous Spaces Using Voronoi Optimistic Optimization with Regret BoundsBeomjoon Kim, Kyungjae Lee, Sungbin Lim, Leslie Pack Kaelbling et al.AAAI 2020 · 55 citations
- Monte-Carlo Tree Search in Continuous Action Spaces with Value GradientsJongmin Lee, Wonseok Jeon, Geon-Hyeong Kim, Kee-Eung KimAAAI 2020 · 24 citations
Related papers
- Improved POMDP Tree Search Planning with Prioritized Action BranchingJohn Mern, Anil Yildiz, Lawrence Bush, Tapan Mukerji et al.AAAI 2021 · 10 citations
- A Surprisingly Simple Continuous-Action POMDP Solver: Lazy Cross-Entropy Search Over Policy TreesMarcus Hörger, Hanna Kurniawati, Dirk P. Kroese, Nan YeAAAI 2024
- Adaptive Online Packing-guided Search for POMDPsChenyang Wu, Guoyu Yang, Zongzhang Zhang, Yang Yu et al.NeurIPS 2021 · 28 citations
- Factored Online Planning in Many-Agent POMDPsMaris F. L. Galesloot, Thiago D. Simão, Sebastian Junges, Nils JansenAAAI 2024 · 3 citations
- POMDP Planning for Object Search in Partially Unknown EnvironmentYongbo Chen, Hanna KurniawatiNeurIPS 2023 · 11 citations
