A Surprisingly Simple Continuous-Action POMDP Solver: Lazy Cross-Entropy Search Over Policy Trees
Marcus Hörger, Hanna Kurniawati, Dirk P. Kroese, Nan Ye
Abstract
The Partially Observable Markov Decision Process (POMDP) provides a principled framework for decision making in stochastic partially observable environments. However, computing good solutions for problems with continuous action spaces remains challenging. To ease this challenge, we propose a simple online POMDP solver, called Lazy Cross-Entropy Search Over Policy Trees (LCEOPT). At each planning step, our method uses a novel lazy Cross-Entropy method to search the space of policy trees, which provide a simple policy representation. Specifically, we maintain a distribution on promising finite-horizon policy trees. The distribution is iteratively updated by sampling policies, evaluating them via Monte Carlo simulation, and refitting them to the top-performing ones. Our method is lazy in the sense that it exploits the policy tree representation to avoid redundant computations in policy sampling, evaluation, and distribution update. This leads to computational savings of up to two orders of magnitude. Our LCEOPT is surprisingly simple as compared to existing state-of-the-art methods, yet empirically outperforms them on several continuous-action POMDP problems, particularly for problems with higher-dimensional action spaces.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on1
Related papers
- Improved POMDP Tree Search Planning with Prioritized Action BranchingJohn Mern, Anil Yildiz, Lawrence Bush, Tapan Mukerji et al.AAAI 2021 · 10 citations
- Adaptive Online Packing-guided Search for POMDPsChenyang Wu, Guoyu Yang, Zongzhang Zhang, Yang Yu et al.NeurIPS 2021 · 28 citations
- POMDP Planning for Object Search in Partially Unknown EnvironmentYongbo Chen, Hanna KurniawatiNeurIPS 2023 · 11 citations
- Information Particle Filter Tree: An Online Algorithm for POMDPs with Belief-Based Rewards on Continuous DomainsJohannes Fischer, Ömer Sahin TasICML 2020 · 42 citations
- Online POMDP Planning with Anytime Deterministic GuaranteesMoran Barenboim, Vadim IndelmanNeurIPS 2023 · 13 citations
