Lune

NeurIPS2024

Worst-Case Offline Reinforcement Learning with Arbitrary Data Support

Kohei Miyaguchi

2024年份

摘要

We propose a method of offline reinforcement learning (RL) featuring the performance guarantee without any assumptions on the data support. Under such conditions, estimating or optimizing the conventional performance metric is generally infeasible due to the distributional discrepancy between data and target policy distributions. To address this issue, we employ a worst-case policy value as a new metric and constructively show that the sample complexity bound of O(ϵ -2 ) is attainable without any data-support conditions, where ϵ > 0 is the policy suboptimality in the new metric. Moreover, as the new metric generalizes the conventional one, the algorithm can address standard offline RL tasks without modification. In this context, our sample complexity bound can be seen as a strict improvement on the previous bounds under the single-policy concentrability and the single-policy realizability. * The author is affiliated with LY Corporation at the time of publication. 38th Conference on Neural Information Processing Systems (NeurIPS 2024). where µ ∈ ∆(S) and β : S → ∆(A) are the behavior state distribution and the behavior policy, respectively. Typically, p M data (D) represents the distribution of the past observational data. The problem of offline RL is now formally given as follows. Problem 3.1 (Offline RL). Given the offline dataset D and a small number ϵ > 0, find a policy π achieving J * -J(π) ≤ ϵ, where J * := max π:S→∆(A) J(π). Value, visitation and weight functions. Let r(s, a) := E y∼R(s,a) [y] be the expected reward function and r π (s) := a r(s, a)π(a|s) be its marginalization with respect to policy π. Let T , T π and