Reference-Based POMDPs
Edward Kim, Yohan Karunanayake, Hanna Kurniawati
摘要
Making good decisions in partially observable and non-deterministic scenarios is a crucial capability for robots. A Partially Observable Markov Decision Process (POMDP) is a general framework for the above problem. Despite advances in POMDP solving, problems with long planning horizons and evolving environments remain difficult to solve even by the best approximate solvers today. To alleviate this difficulty, we propose a slightly modified POMDP problem, called a Reference-Based POMDP, where the objective is to balance between maximizing the expected total reward and being close to a given reference (stochastic) policy. The optimal policy of a Reference-Based POMDP can be computed via iterative expectations using the given reference policy, thereby avoiding exhaustive enumeration of actions at each belief node of the search tree. We demonstrate theoretically that the standard POMDP under stochastic policies is related to the Reference-Based POMDP. To demonstrate the feasibility of exploiting the formulation, we present a basic algorithm R EF S OLVER . Results from experiments on long-horizon navigation problems indicate that this basic algorithm substantially outperforms POMCP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Search and Explore: Symbiotic Policy Synthesis in POMDPsRoman Andriushchenko, Alexander Bork, Milan Ceska, Sebastian Junges 等CAV 2023 · 被引用 7 次
- Point-Based Methods for Model Checking in Partially Observable Markov Decision ProcessesMaxime Bouton, Jana Tumova, Mykel J. KochenderferAAAI 2020 · 被引用 32 次
- The Geometry of Memoryless Stochastic Policy Optimization in Infinite-Horizon POMDPsJohannes Müller, Guido MontúfarICLR 2022 · 被引用 9 次
- Information Particle Filter Tree: An Online Algorithm for POMDPs with Belief-Based Rewards on Continuous DomainsJohannes Fischer, Ömer Sahin TasICML 2020 · 被引用 42 次
- Scalable Policy-Based RL Algorithms for POMDPsAmeya Anjarlekar, S. Rasoul Etesami, R. SrikantNeurIPS 2025 · 被引用 6 次
