Lune

NeurIPS2023顶会

Provably Safe Reinforcement Learning with Step-wise Violation Constraints

Nuoya Xiong, Yihan Du, Longbo Huang

2023年份
16被引次数
6顶会引用

摘要

In this paper, we investigate a novel safe reinforcement learning problem with step-wise violation constraints. Our problem differs from existing works in that we consider stricter step-wise violation constraints and do not assume the existence of safe actions, making our formulation more suitable for safety-critical applications which need to ensure safety in all decision steps and may not always possess safe actions, e.g., robot control and autonomous driving. We propose a novel algorithm SUCBVI, which guarantees O~(ST)\widetilde{O}(\sqrt{ST}) step-wise violation and O~(H3SAT)\widetilde{O}(\sqrt{H^3SAT}) regret. Lower bounds are provided to validate the optimality in both violation and regret performance with respect to SS and TT. Moreover, we further study a novel safe reward-free exploration problem with step-wise violation constraints. For this problem, we design an (ε,δ)(\varepsilon,\delta)-PAC algorithm SRF-UCRL, which achieves nearly state-of-the-art sample complexity O~((S2AH2ε+H4SAε2)(log⁡(1δ)+S))\widetilde{O}((\frac{S^2AH^2}{\varepsilon}+\frac{H^4SA}{\varepsilon^2})(\log(\frac{1}{\delta})+S)), and guarantees O~(ST)\widetilde{O}(\sqrt{ST}) violation during the exploration. The experimental results demonstrate the superiority of our algorithms in safety performance, and corroborate our theoretical results.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper6

问问它们各自怎么用它

它引用的顶会 Paper19

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖