Online POMDP Planning with Anytime Deterministic Guarantees
Moran Barenboim, Vadim Indelman
摘要
Decision-making under uncertainty is a critical aspect of many practical autonomous systems due to incomplete information. Partially Observable Markov Decision Processes (POMDPs) offer a mathematically principled framework for formulating decision-making problems under such conditions. However, finding an optimal solution for a POMDP is generally intractable. In recent years, there has been a significant progress of scaling approximate solvers from small to moderately sized problems, using online tree search solvers. Often, such approximate solvers are limited to probabilistic or asymptotic guarantees towards the optimal solution. In this paper, we derive a deterministic relationship for discrete POMDPs between an approximated and the optimal solution. We show that at any time, we can derive bounds that relate between the existing solution and the optimal one. We show that our derivations provide an avenue for a new set of algorithms and can be attached to existing algorithms that have a certain structure to provide them with deterministic guarantees with marginal computational overhead. In return, not only do we certify the solution quality, but we demonstrate that making a decision based on the deterministic guarantee may result in superior performance compared to the original algorithm without the deterministic certification. Introduction Decision-making under uncertainty is a common challenge in many practical autonomous systems. In such systems, agents often operate with incomplete information about their environment. This uncertainty can arise from various sources, including sensor noise, hardware limitations, modeling approximations, 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Belief-State Query Policies for User-Aligned POMDPsDaniel Bramblett, Siddharth SrivastavaNeurIPS 2024 · 被引用 1 次
- Counterfactual Online Learning for Open-Loop Monte-Carlo PlanningThomy Phan, Shao-Hung Chan, Sven KoenigAAAI 2025
相关 Paper
- Scalable Policy-Based RL Algorithms for POMDPsAmeya Anjarlekar, S. Rasoul Etesami, R. SrikantNeurIPS 2025 · 被引用 6 次
- Efficient Multiagent Planning via Shared Action SuggestionsDylan M. Asmar, Mykel J. KochenderferAAAI 2026
- Revelations: A Decidable Class of POMDPs with Omega-Regular ObjectivesMarius Belly, Nathanaël Fijalkow, Hugo Gimbert, Florian Horn 等AAAI 2025 · 被引用 5 次
- A Surprisingly Simple Continuous-Action POMDP Solver: Lazy Cross-Entropy Search Over Policy TreesMarcus Hörger, Hanna Kurniawati, Dirk P. Kroese, Nan YeAAAI 2024
- Revealing POMDPs: Qualitative and Quantitative Analysis for Parity ObjectivesAli Asadi, Krishnendu Chatterjee, David Lurie, Raimundo SaonaAAAI 2026 · 被引用 1 次
