Lune

ICLR2025顶会

Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form

Toshinori Kitamura, Tadashi Kozuno, Wataru Kumagai, Kenta Hoshino, Yohei Hosoe, Kazumi Kasaura, Masashi Hamaya, Paavo Parmas, Yutaka Matsuo

出版方
2025年份
2顶会引用

摘要

Erratum. We draw the reader's attention to a technical error in the proof of global optimality of stationary points of the maximum violation function ∆ b0 (π) in Theorem 4. Specifically, Equation ( 37 ) is false in general. The proof requires the identity min y∈Y max x∈X ⟨x, y⟩ = min y∈conv(Y) max x∈X ⟨x, y⟩ for compact sets X , Y ⊆ R d with X convex, but this identity does not hold in general; for instance, one may take Kitamura et al. (2026) for details. Unfortunately, Kitamura et al. (2026) further show that general RMDPs may admit multiple local minima. Moreover, even under (s, a)-rectangularity, finding an ε-optimal policy in RCMDPs is NP-hard. As a result, the main result of this paper, Corollary 1, is also incorrect.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper18

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖