Robust -Divergence MDPs
Chin Pang Ho, Marek Petrik, Wolfram Wiesemann
摘要
In recent years, robust Markov decision processes (MDPs) have emerged as a prominent modeling framework for dynamic decision problems affected by uncertainty. In contrast to classical MDPs, which only account for stochasticity by modeling the dynamics through a stochastic process with a known transition kernel, robust MDPs additionally account for ambiguity by optimizing in view of the most adverse transition kernel from a prescribed ambiguity set. In this paper, we develop a novel solution framework for robust MDPs with s-rectangular ambiguity sets that decomposes the problem into a sequence of robust Bellman updates and simplex projections. Exploiting the rich structure present in the simplex projections corresponding to φ-divergence ambiguity sets, we show that the associated s-rectangular robust MDPs can be solved substantially faster than with state-of-the-art commercial solvers as well as a recent first-order solution scheme, thus rendering them attractive alternatives to classical MDPs in practical applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Adjustable Robust Reinforcement Learning for Online 3D Bin PackingYuxin Pan, Yize Chen, Fangzhen LinNeurIPS 2023 · 被引用 23 次
- Fast Bellman Updates for Wasserstein Distributionally Robust MDPsZhuodong Yu, Ling Dai, Shaohang Xu, Siyang Gao 等NeurIPS 2023 · 被引用 15 次
- Non-rectangular Robust MDPs with Normed Uncertainty SetsNavdeep Kumar, Adarsh Gupta, Maxence Mohamed Elfatihi, Giorgia Ramponi 等NeurIPS 2025 · 被引用 3 次
- Robust Satisficing MDPsHaolin Ruan, Siyu Zhou, Zhi Chen, Chin Pang HoICML 2023 · 被引用 2 次
- Convergence of Fast Policy Iteration in Markov Games and Robust MDPsKeith Badger, Jefferson Huang, Marek PetrikAAAI 2026
它引用的顶会 Paper9
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 被引用 157 次
- Policy Gradient Method For Robust Reinforcement LearningYue Wang, Shaofeng ZouICML 2022 · 被引用 104 次
- Robust Reinforcement Learning using Least Squares Policy Iteration with Provable Performance GuaranteesKishan Panaganti Badrinath, Dileep KalathilICML 2021 · 被引用 78 次
- Twice regularized MDPs and the equivalence between robustness and regularizationEsther Derman, Matthieu Geist, Shie MannorNeurIPS 2021 · 被引用 68 次
- Navigating to the Best Policy in Markov Decision ProcessesAymen Al Marjani, Aurélien Garivier, Alexandre ProutièreNeurIPS 2021 · 被引用 34 次
相关 Paper
- Fast Algorithms for -constrained S-rectangular Robust MDPsBahram Behzadian, Marek Petrik, Chin Pang HoNeurIPS 2021 · 被引用 4 次
- Solving Robust Markov Decision Processes: Generic, Reliable, EfficientTobias Meggendorfer, Maximilian Weininger, Patrick WienhöftAAAI 2025
- Scalable First-Order Methods for Robust MDPsJulien Grand-Clément, Christian KroerAAAI 2021 · 被引用 33 次
- Provable Policy Gradient for Robust Average-Reward MDPs Beyond RectangularityQiuhao Wang, Yuqi Zha, Chin Pang Ho, Marek PetrikICML 2025
- Efficient Value Iteration for s-rectangular Robust Markov Decision ProcessesNavdeep Kumar, Kaixin Wang, Kfir Yehuda Levy, Shie MannorICML 2024 · 被引用 4 次
