Robust Satisficing MDPs
Haolin Ruan, Siyu Zhou, Zhi Chen, Chin Pang Ho
Abstract
Despite being a fundamental building block for reinforcement learning, Markov decision processes (MDPs) often suffer from ambiguity in model parameters. Robust MDPs are proposed to overcome this challenge by optimizing the worstcase performance under ambiguity. While robust MDPs can provide reliable policies with limited data, their worst-case performances are often overly conservative, and so they do not offer practical insights into the actual performance of these reliable policies. This paper proposes robust satisficing MDPs (RSMDPs), where the expected returns of feasible policies are softlyconstrained to achieve a user-specified target under ambiguity. We derive a tractable reformulation for RSMDPs and develop a first-order method for solving large instances. Experimental results demonstrate that RSMDPs can prescribe policies to achieve their targets, which are much higher than the optimal worst-case returns computed by robust MDPs. Moreover, the average and percentile performances of our model are competitive among other models. We also demonstrate the scalability of the proposed algorithm compared with a state-of-the-art commercial solver. Introduction Markov decision processes (MDPs) have emerged as a powerful modeling framework for sequential decision-making problems under uncertainty (Ashok et al., 2019; Puterman, 2014; Sutton & Barto, 2018) . Successful employments of MDPs largely rely on the perfect estimation of model parameters (Petrik & Russel, 2019), which, unfortunately, is not always the case. A common situation is when the true parameters are estimated from a limited amount of sam-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ce665963-c716-4610-ae67-2311c5b20c8aCited by top-tier papers2
- Iterative Robust Satisficing: Minimizing Performance Degradation Under Distribution ShiftEnes Ağırman, Artun Saday, Cem TekinICML 2026
- Statistical Properties of Robust SatisficingZhiyi Li, Yunbei Xu, Ruohan ZhanICML 2024
Builds on8
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 157 citations
- Policy Gradient Method For Robust Reinforcement LearningYue Wang, Shaofeng ZouICML 2022 · 104 citations
- Twice regularized MDPs and the equivalence between robustness and regularizationEsther Derman, Matthieu Geist, Shie MannorNeurIPS 2021 · 68 citations
- Bayesian Robust Optimization for Imitation LearningDaniel S. Brown, Scott Niekum, Marek PetrikNeurIPS 2020 · 43 citations
- Scalable First-Order Methods for Robust MDPsJulien Grand-Clément, Christian KroerAAAI 2021 · 33 citations
Related papers
- Robust -Divergence MDPsChin Pang Ho, Marek Petrik, Wolfram WiesemannNeurIPS 2022 · 13 citations
- Provable Policy Gradient for Robust Average-Reward MDPs Beyond RectangularityQiuhao Wang, Yuqi Zha, Chin Pang Ho, Marek PetrikICML 2025
- First-Order Methods for Wasserstein Distributionally Robust MDPJulien Grand-Clément, Christian KroerICML 2021 · 32 citations
- Fast Algorithms for -constrained S-rectangular Robust MDPsBahram Behzadian, Marek Petrik, Chin Pang HoNeurIPS 2021 · 4 citations
- Constrained and Robust Policy Synthesis with Satisfiability-Modulo-Probabilistic-Model-CheckingLinus Heck, Filip Macák, Milan Ceska, Sebastian JungesAAAI 2026
