Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
Toshinori Kitamura, Tadashi Kozuno, Wataru Kumagai, Kenta Hoshino, Yohei Hosoe, Kazumi Kasaura, Masashi Hamaya, Paavo Parmas, Yutaka Matsuo
摘要
Erratum. We draw the reader's attention to a technical error in the proof of global optimality of stationary points of the maximum violation function ∆ b0 (π) in Theorem 4. Specifically, Equation ( 37 ) is false in general. The proof requires the identity min y∈Y max x∈X ⟨x, y⟩ = min y∈conv(Y) max x∈X ⟨x, y⟩ for compact sets X , Y ⊆ R d with X convex, but this identity does not hold in general; for instance, one may take Kitamura et al. (2026) for details. Unfortunately, Kitamura et al. (2026) further show that general RMDPs may admit multiple local minima. Moreover, even under (s, a)-rectangularity, finding an ε-optimal policy in RCMDPs is NP-hard. As a result, the main result of this paper, Corollary 1, is also incorrect.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SPiDR: A Simple Approach for Zero-Shot Safety in Sim-to-Real TransferYarden As, Chengrui Qu, Benjamin Unger, Dongho Kang 等NeurIPS 2025 · 被引用 9 次
- Efficient Policy Optimization in Robust Constrained MDPs with Iteration Complexity GuaranteesSourav Ganguly, Kishan Panaganti, Arnob Ghosh, Adam WiermanNeurIPS 2025 · 被引用 7 次
它引用的顶会 Paper18
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 被引用 252 次
- Bilinear Classes: A Structural Framework for Provable Generalization in RLSimon S. Du, Sham M. Kakade, Jason D. Lee, Shachar Lovett 等ICML 2021 · 被引用 207 次
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 被引用 157 次
- Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPsTao Liu, Ruida Zhou, Dileep Kalathil, Panganamala R. Kumar 等NeurIPS 2021 · 被引用 110 次
相关 Paper
- Fast Algorithms for -constrained S-rectangular Robust MDPsBahram Behzadian, Marek Petrik, Chin Pang HoNeurIPS 2021 · 被引用 4 次
- A Single-Loop Robust Policy Gradient Method for Robust Markov Decision ProcessesZhenwei Lin, Chenyu Xue, Qi Deng, Yinyu YeICML 2024 · 被引用 3 次
- Optimal Strong Regret and Violation in Constrained MDPs via Policy OptimizationFrancesco Emanuele Stradi, Matteo Castiglioni, Alberto Marchesi, Nicola GattiICLR 2025
- Achieving Zero Constraint Violation for Constrained Reinforcement Learning via Conservative Natural Policy Gradient Primal-Dual AlgorithmQinbo Bai, Amrit Singh Bedi, Vaneet AggarwalAAAI 2023 · 被引用 29 次
- Learning General Parameterized Policies for Infinite Horizon Average Reward Constrained MDPs via Primal-Dual Policy Gradient AlgorithmQinbo Bai, Washim Uddin Mondal, Vaneet AggarwalNeurIPS 2024 · 被引用 10 次
