Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
Toshinori Kitamura, Tadashi Kozuno, Wataru Kumagai, Kenta Hoshino, Yohei Hosoe, Kazumi Kasaura, Masashi Hamaya, Paavo Parmas, Yutaka Matsuo
Abstract
Erratum. We draw the reader's attention to a technical error in the proof of global optimality of stationary points of the maximum violation function ∆ b0 (π) in Theorem 4. Specifically, Equation ( 37 ) is false in general. The proof requires the identity min y∈Y max x∈X ⟨x, y⟩ = min y∈conv(Y) max x∈X ⟨x, y⟩ for compact sets X , Y ⊆ R d with X convex, but this identity does not hold in general; for instance, one may take Kitamura et al. (2026) for details. Unfortunately, Kitamura et al. (2026) further show that general RMDPs may admit multiple local minima. Moreover, even under (s, a)-rectangularity, finding an ε-optimal policy in RCMDPs is NP-hard. As a result, the main result of this paper, Corollary 1, is also incorrect.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d95b1f75-d1a4-45af-a603-803bb1f83d61Cited by top-tier papers2
- SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real TransferYarden As, Chengrui Qu, Benjamin Unger, Dongho Kang et al.NeurIPS 2025 · 9 citations
- Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity GuaranteesSourav Ganguly, Kishan Panaganti, Arnob Ghosh, Adam WiermanNeurIPS 2025 · 7 citations
Builds on18
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 252 citations
- Bilinear Classes: A Structural Framework for Provable Generalization in RLSimon S. Du, Sham M. Kakade, Jason D. Lee, Shachar Lovett et al.ICML 2021 · 207 citations
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 157 citations
- Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPsTao Liu, Ruida Zhou, Dileep Kalathil, Panganamala R. Kumar et al.NeurIPS 2021 · 110 citations
Related papers
- Fast Algorithms for -constrained S-rectangular Robust MDPsBahram Behzadian, Marek Petrik, Chin Pang HoNeurIPS 2021 · 4 citations
- A Single-Loop Robust Policy Gradient Method for Robust Markov Decision ProcessesZhenwei Lin, Chenyu Xue, Qi Deng, Yinyu YeICML 2024 · 3 citations
- Optimal Strong Regret and Violation in Constrained MDPs via Policy OptimizationFrancesco Emanuele Stradi, Matteo Castiglioni, Alberto Marchesi, Nicola GattiICLR 2025
- Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual AlgorithmQinbo Bai, Amrit Singh Bedi, Vaneet AggarwalAAAI 2023 · 29 citations
- Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient AlgorithmQinbo Bai, Washim Uddin Mondal, Vaneet AggarwalNeurIPS 2024 · 10 citations
