Bayesian Optimized Monte Carlo Planning
John Mern, Anil Yildiz, Zachary Sunberg, Tapan Mukerji, Mykel J. Kochenderfer
摘要
Online solvers for partially observable Markov decision processes have difficulty scaling to problems with large action spaces. Monte Carlo tree search with progressive widening attempts to improve scaling by sampling from the action space to construct a policy search tree. The performance of progressive widening search is dependent upon the action sampling policy, often requiring problem-specific samplers. In this work, we present a general method for efficient action sampling based on Bayesian optimization. The proposed method uses a Gaussian process to model a belief over the action-value function and selects the action that will maximize the expected improvement in the optimal action value. We implement the proposed approach in a new online tree search algorithm called Bayesian Optimized Monte Carlo Planning (BOMCP). Several experiments show that BOMCP is better able to scale to large action space POMDPs than existing state-of-the-art tree search solvers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Risk-Averse Bayes-Adaptive Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2021 · 被引用 50 次
- Information-guided Planning: An Online Approach for Partially Observable ProblemsMatheus Aparecido do Carmo Alves, Amokh Varma, Yehia Elkhatib, Leandro Soriano MarcolinoNeurIPS 2023 · 被引用 2 次
- A Bayesian Approach to Online PlanningNir Greshler, David Ben-Eli, Carmel Rabinovitz, Gabi Guetta 等ICML 2024 · 被引用 1 次
- ReqsMiner: Automated Discovery of CDN Forwarding Request Inconsistencies and DoS Attacks with Grammar-based FuzzingLinkai Zheng, Xiang Li, Chuhan Wang, Run Guo 等NDSS 2024
- Epistemic Monte Carlo Tree SearchYaniv Oren, Viliam Vadocz, Matthijs T. J. Spaan, Wendelin BoehmerICLR 2025
它引用的顶会 Paper2
- Monte Carlo Tree Search in Continuous Spaces Using Voronoi Optimistic Optimization with Regret BoundsBeomjoon Kim, Kyungjae Lee, Sungbin Lim, Leslie Pack Kaelbling 等AAAI 2020 · 被引用 55 次
- Monte-Carlo Tree Search in Continuous Action Spaces with Value GradientsJongmin Lee, Wonseok Jeon, Geon-Hyeong Kim, Kee-Eung KimAAAI 2020 · 被引用 24 次
相关 Paper
- Improved POMDP Tree Search Planning with Prioritized Action BranchingJohn Mern, Anil Yildiz, Lawrence Bush, Tapan Mukerji 等AAAI 2021 · 被引用 10 次
- A Surprisingly Simple Continuous-Action POMDP Solver: Lazy Cross-Entropy Search Over Policy TreesMarcus Hörger, Hanna Kurniawati, Dirk P. Kroese, Nan YeAAAI 2024
- Adaptive Online Packing-guided Search for POMDPsChenyang Wu, Guoyu Yang, Zongzhang Zhang, Yang Yu 等NeurIPS 2021 · 被引用 28 次
- Factored Online Planning in Many-Agent POMDPsMaris F. L. Galesloot, Thiago D. Simão, Sebastian Junges, Nils JansenAAAI 2024 · 被引用 3 次
- POMDP Planning for Object Search in Partially Unknown EnvironmentYongbo Chen, Hanna KurniawatiNeurIPS 2023 · 被引用 11 次
