Robust Anytime Learning of Markov Decision Processes
Marnix Suilen, Thiago D. Simão, David Parker, Nils Jansen
Abstract
Markov decision processes (MDPs) are formal models commonly used in sequential decision-making. MDPs capture the stochasticity that may arise, for instance, from imprecise actuators via probabilities in the transition function. However, in data-driven applications, deriving precise probabilities from (limited) data introduces statistical errors that may lead to unexpected or undesirable outcomes. Uncertain MDPs (uMDPs) do not require precise probabilities but instead use so-called uncertainty sets in the transitions, accounting for such limited data. Tools from the formal verification community efficiently compute robust policies that provably adhere to formal specifications, like safety constraints, under the worst-case instance in the uncertainty set. We continuously learn the transition probabilities of an MDP in a robust anytime-learning approach that combines a dedicated Bayesian inference scheme with the computation of robust policies. In particular, our method (1) approximates probabilities as intervals, (2) adapts to new data that may be inconsistent with an intermediate model, and (3) may be stopped at any time to compute a robust policy on the uMDP that faithfully captures the data so far. Furthermore, our method is capable of adapting to changes in the environment. We show the effectiveness of our approach and compare it to robust policies computed on uMDPs learned by the UCRL2 reinforcement learning algorithm in an experimental evaluation on several benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 358891fd-70aa-4bb1-afa2-9cd8a2b3ed92Cited by top-tier papers5
- On Evaluating Policies for Robust POMDPsMerlijn Krale, Eline M. Bovy, Maris F. L. Galesloot, Thiago D. Simão et al.NeurIPS 2025 · 2 citations
- Robust Satisficing MDPsHaolin Ruan, Siyu Zhou, Zhi Chen, Chin Pang HoICML 2023 · 2 citations
- Efficient Solution and Learning of Robust Factored MDPsYannik Schnitzer, Alessandro Abate, David ParkerAAAI 2026 · 1 citation
- Robust Transfer of Safety-Constrained Reinforcement Learning AgentsMarkel Zubia, Thiago D. Simão, Nils JansenICLR 2025
- Online Robust Planning Under Model Uncertainty: A Sample-Based ApproachTamir Shazman, Idan Lev-Yehudi, Ron Benchetrit, Vadim IndelmanAAAI 2026
Builds on6
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 419 citations
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao et al.NeurIPS 2021 · 373 citations
- Learning Adversarial Markov Decision Processes with Bandit Feedback and Unknown TransitionChi Jin, Tiancheng Jin, Haipeng Luo, Suvrit Sra et al.ICML 2020 · 117 citations
- Robust Finite-State Controllers for Uncertain POMDPsMurat Cubuktepe, Nils Jansen, Sebastian Junges, Ahmadreza Marandi et al.AAAI 2021 · 35 citations
- The Importance of Pessimism in Fixed-Dataset Policy OptimizationJacob Buckman, Carles Gelada, Marc G. BellemareICLR 2021 · 23 citations
Related papers
- Solving Robust Markov Decision Processes: Generic, Reliable, EfficientTobias Meggendorfer, Maximilian Weininger, Patrick WienhöftAAAI 2025
- Bayesian Risk-Averse Q-Learning with Streaming ObservationsYuhao Wang, Enlu ZhouNeurIPS 2023 · 8 citations
- Approximate Bilevel Difference Convex Programming for Bayesian Risk Markov Decision ProcessesYifan Lin, Enlu ZhouAAAI 2025 · 1 citation
- Online MDP with Prototypes Information: A Robust Adaptive ApproachShuo Sun, Meng Qi, Zuo-Jun Max ShenAAAI 2025 · 2 citations
- Model-Free Robust Average-Reward Reinforcement LearningYue Wang, Alvaro Velasquez, George K. Atia, Ashley Prater-Bennette et al.ICML 2023 · 25 citations
