Lune

NeurIPS2021顶会

Optimal Uniform OPE and Model-based Offline Reinforcement Learning in Time-Homogeneous, Reward-Free and Task-Agnostic Settings

Ming Yin, Yu-Xiang Wang

2021年份
19被引次数
10顶会引用

摘要

This work studies the statistical limits of uniform convergence for offline policy evaluation (OPE) problems with model-based methods (for episodic MDP) and provides a unified framework towards optimal learning for several well-motivated offline tasks. Uniform OPE sup⁡Π∣Qπ−Q^π∣<ϵ\sup_\Pi|Q^\pi-\hat{Q}^\pi|<\epsilon is a stronger measure than the point-wise OPE and ensures offline learning when Π\Pi contains all policies (the global class). In this paper, we establish an Ω(H2S/dmϵ2)\Omega(H^2 S/d_m\epsilon^2) lower bound (over model-based family) for the global uniform OPE and our main result establishes an upper bound of O~(H2/dmϵ2)\tilde{O}(H^2/d_m\epsilon^2) for the local uniform convergence that applies to all near-empirically optimal policies for the MDPs with stationary transition. Here dmd_m is the minimal marginal state-action probability. Critically, the highlight in achieving the optimal rate O~(H2/dmϵ2)\tilde{O}(H^2/d_m\epsilon^2) is our design of singleton absorbing MDP, which is a new sharp analysis tool that works with the model-based approach. We generalize such a model-based framework to the new settings: offline task-agnostic and the offline reward-free with optimal complexity O~(H2log⁡(K)/dmϵ2)\tilde{O}(H^2\log(K)/d_m\epsilon^2) (KK is the number of tasks) and O~(H2S/dmϵ2)\tilde{O}(H^2S/d_m\epsilon^2) respectively. These results provide a unified solution for simultaneously solving different offline RL problems.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper10

问问它们各自怎么用它

它引用的顶会 Paper21

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖