On Evaluating Policies for Robust POMDPs
Merlijn Krale, Eline M. Bovy, Maris F. L. Galesloot, Thiago D. Simão, Nils Jansen
Abstract
Robust partially observable Markov decision processes (RPOMDPs) model sequential decision-making problems under partial observability, where an agent must be robust against a range of dynamics. RPOMDPs can be viewed as a two-player game between an agent, who selects actions, and nature , who adversarially selects the dynamics. Evaluating an agent policy requires finding an adversarial nature policy, which is computationally challenging. In this paper, we advance the evaluation of agent policies for RPOMDPs in three ways. First, we discuss suitable benchmarks. We observe that for some RPOMDPs, an optimal agent policy can be found by considering only subsets of nature policies, making them easier to solve. We formalize this concept of solvability and construct three benchmarks that are only solvable for expressive sets of nature policies. Second, we describe a new method to evaluate agent policies for RPOMDPs by solving an equivalent MDP. Third, we lift two well-known upper bounds from POMDPs to RPOMDPs, which can be used to efficiently approximate the optimality gap of a policy and serve as baselines. Our experimental evaluation shows that (1) our proposed benchmarks cannot be solved by assuming naive nature policies, (2) our method of evaluating policies is accurate, and (3) the upper bounds provide solid baselines for evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 566015e9-c810-4fa2-ace4-9b33337e0230Builds on6
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides et al.ICLR 2020 · 204 citations
- Robust Finite-State Controllers for Uncertain POMDPsMurat Cubuktepe, Nils Jansen, Sebastian Junges, Ahmadreza Marandi et al.AAAI 2021 · 35 citations
- Sampling-Based Robust Control of Autonomous Systems with Non-Gaussian NoiseThom S. Badings, Alessandro Abate, Nils Jansen, David Parker et al.AAAI 2022 · 33 citations
- Robust Anytime Learning of Markov Decision ProcessesMarnix Suilen, Thiago D. Simão, David Parker, Nils JansenNeurIPS 2022 · 31 citations
- Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement LearningShangding Gu, Laixi Shi, Muning Wen, Ming Jin et al.ICLR 2025
Related papers
- Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial ObservabilityEline M. Bovy, Caleb Probine, Marnix Suilen, Ufuk Topcu et al.NeurIPS 2025 · 3 citations
- Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount FactorJulien Grand-Clément, Marek PetrikNeurIPS 2023 · 25 citations
- Qualitative Analysis of ω-Regular Objectives on Robust MDPsAli Asadi, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Mehrdad Karrabi et al.AAAI 2026
- Revealing POMDPs: Qualitative and Quantitative Analysis for Parity ObjectivesAli Asadi, Krishnendu Chatterjee, David Lurie, Raimundo SaonaAAAI 2026 · 1 citation
- Enforcing Almost-Sure Reachability in POMDPsSebastian Junges, Nils Jansen, Sanjit A. SeshiaCAV 2021 · 8 citations
