Robust Satisficing MDPs
Haolin Ruan, Siyu Zhou, Zhi Chen, Chin Pang Ho
摘要
Despite being a fundamental building block for reinforcement learning, Markov decision processes (MDPs) often suffer from ambiguity in model parameters. Robust MDPs are proposed to overcome this challenge by optimizing the worstcase performance under ambiguity. While robust MDPs can provide reliable policies with limited data, their worst-case performances are often overly conservative, and so they do not offer practical insights into the actual performance of these reliable policies. This paper proposes robust satisficing MDPs (RSMDPs), where the expected returns of feasible policies are softlyconstrained to achieve a user-specified target under ambiguity. We derive a tractable reformulation for RSMDPs and develop a first-order method for solving large instances. Experimental results demonstrate that RSMDPs can prescribe policies to achieve their targets, which are much higher than the optimal worst-case returns computed by robust MDPs. Moreover, the average and percentile performances of our model are competitive among other models. We also demonstrate the scalability of the proposed algorithm compared with a state-of-the-art commercial solver. Introduction Markov decision processes (MDPs) have emerged as a powerful modeling framework for sequential decision-making problems under uncertainty (Ashok et al., 2019; Puterman, 2014; Sutton & Barto, 2018) . Successful employments of MDPs largely rely on the perfect estimation of model parameters (Petrik & Russel, 2019), which, unfortunately, is not always the case. A common situation is when the true parameters are estimated from a limited amount of sam-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Iterative Robust Satisficing: Minimizing Performance Degradation Under Distribution ShiftEnes Ağırman, Artun Saday, Cem TekinICML 2026
- Statistical Properties of Robust SatisficingZhiyi Li, Yunbei Xu, Ruohan ZhanICML 2024
它引用的顶会 Paper8
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 被引用 157 次
- Policy Gradient Method For Robust Reinforcement LearningYue Wang, Shaofeng ZouICML 2022 · 被引用 104 次
- Twice regularized MDPs and the equivalence between robustness and regularizationEsther Derman, Matthieu Geist, Shie MannorNeurIPS 2021 · 被引用 68 次
- Bayesian Robust Optimization for Imitation LearningDaniel S. Brown, Scott Niekum, Marek PetrikNeurIPS 2020 · 被引用 43 次
- Scalable First-Order Methods for Robust MDPsJulien Grand-Clément, Christian KroerAAAI 2021 · 被引用 33 次
相关 Paper
- Robust -Divergence MDPsChin Pang Ho, Marek Petrik, Wolfram WiesemannNeurIPS 2022 · 被引用 13 次
- Provable Policy Gradient for Robust Average-Reward MDPs Beyond RectangularityQiuhao Wang, Yuqi Zha, Chin Pang Ho, Marek PetrikICML 2025
- First-Order Methods for Wasserstein Distributionally Robust MDPJulien Grand-Clément, Christian KroerICML 2021 · 被引用 32 次
- Fast Algorithms for -constrained S-rectangular Robust MDPsBahram Behzadian, Marek Petrik, Chin Pang HoNeurIPS 2021 · 被引用 4 次
- Constrained and Robust Policy Synthesis with Satisfiability-Modulo-Probabilistic-Model-CheckingLinus Heck, Filip Macák, Milan Ceska, Sebastian JungesAAAI 2026
