On Evaluating Policies for Robust POMDPs
Merlijn Krale, Eline M. Bovy, Maris F. L. Galesloot, Thiago D. Simão, Nils Jansen
摘要
Robust partially observable Markov decision processes (RPOMDPs) model sequential decision-making problems under partial observability, where an agent must be robust against a range of dynamics. RPOMDPs can be viewed as a two-player game between an agent, who selects actions, and nature , who adversarially selects the dynamics. Evaluating an agent policy requires finding an adversarial nature policy, which is computationally challenging. In this paper, we advance the evaluation of agent policies for RPOMDPs in three ways. First, we discuss suitable benchmarks. We observe that for some RPOMDPs, an optimal agent policy can be found by considering only subsets of nature policies, making them easier to solve. We formalize this concept of solvability and construct three benchmarks that are only solvable for expressive sets of nature policies. Second, we describe a new method to evaluate agent policies for RPOMDPs by solving an equivalent MDP. Third, we lift two well-known upper bounds from POMDPs to RPOMDPs, which can be used to efficiently approximate the optimality gap of a policy and serve as baselines. Our experimental evaluation shows that (1) our proposed benchmarks cannot be solved by assuming naive nature policies, (2) our method of evaluating policies is accurate, and (3) the upper bounds provide solid baselines for evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides 等ICLR 2020 · 被引用 204 次
- Robust Finite-State Controllers for Uncertain POMDPsMurat Cubuktepe, Nils Jansen, Sebastian Junges, Ahmadreza Marandi 等AAAI 2021 · 被引用 35 次
- Sampling-Based Robust Control of Autonomous Systems with Non-Gaussian NoiseThom S. Badings, Alessandro Abate, Nils Jansen, David Parker 等AAAI 2022 · 被引用 33 次
- Robust Anytime Learning of Markov Decision ProcessesMarnix Suilen, Thiago D. Simão, David Parker, Nils JansenNeurIPS 2022 · 被引用 31 次
- Robust Gymnasium: A Unified Modular Benchmark for Robust Reinforcement LearningShangding Gu, Laixi Shi, Muning Wen, Ming Jin 等ICLR 2025
相关 Paper
- Multi-Environment POMDPs: Discrete Model Uncertainty Under Partial ObservabilityEline M. Bovy, Caleb Probine, Marnix Suilen, Ufuk Topcu 等NeurIPS 2025 · 被引用 3 次
- Reducing Blackwell and Average Optimality to Discounted MDPs via the Blackwell Discount FactorJulien Grand-Clément, Marek PetrikNeurIPS 2023 · 被引用 25 次
- Qualitative Analysis of ω-Regular Objectives on Robust MDPsAli Asadi, Krishnendu Chatterjee, Ehsan Kafshdar Goharshady, Mehrdad Karrabi 等AAAI 2026
- Revealing POMDPs: Qualitative and Quantitative Analysis for Parity ObjectivesAli Asadi, Krishnendu Chatterjee, David Lurie, Raimundo SaonaAAAI 2026 · 被引用 1 次
- Enforcing Almost-Sure Reachability in POMDPsSebastian Junges, Nils Jansen, Sanjit A. SeshiaCAV 2021 · 被引用 8 次
