Lune

NeurIPS2023顶会

Double Pessimism is Provably Efficient for Distributionally Robust Offline Reinforcement Learning: Generic Algorithm and Robust Partial Coverage

Jose H. Blanchet, Miao Lu, Tong Zhang, Han Zhong

2023年份
58被引次数
34顶会引用

摘要

In this paper, we study distributionally robust offline reinforcement learning (robust offline RL), which seeks to find an optimal policy purely from an offline dataset that can perform well in perturbed environments. In specific, we propose a generic algorithm framework called Doubly Pessimistic Model-based Policy Optimization (P 2 MPO), which features a novel combination of a flexible model estimation subroutine and a doubly pessimistic policy optimization step. Notably, the double pessimism principle is crucial to overcome the distributional shifts incurred by (i) the mismatch between the behavior policy and the family of target policies; and (ii) the perturbation of the nominal model. Under certain accuracy conditions on the model estimation subroutine, we prove that P 2 MPO is sample-efficient with robust partial coverage data, which only requires the offline data to have good coverage of the distributions induced by the optimal robust policy and the perturbed models around the nominal model. Our assumption on data is relatively mild compared with previous full-coverage-style assumptions which need a uniformly lower bounded data distribution. Our algorithm and theory can be applied to a vast body of robust Markov decision processes (RMDPs) in the regime of large state spaces. By tailoring specific model estimation subroutines for concrete examples of RMDPs, including tabular RMDPs, factored RMDPs, kernel and neural RMDPs, we prove that for all these examples P 2 MPO enjoys a O(n -1/2 ) convergence rate, where n is the number of trajectories in data. We highlight that all these RMDP examples, except tabular RMDPs, are first identified and proven tractable by this work. Furthermore, as an extension to multi-agent decision-making, we continue our study of robust offline RL in the multi-player robust Markov games (RMGs). By extending the double pessimism principle identified for single-agent RMDPs, we propose another doubly-pessimistic-type algorithm framework that can efficiently find the robust Nash equilibria among players using only robust unilateral (partial) coverage data. To our best knowledge, this work proposes the first general learning principle -double pessimismfor robust offline RL and shows that it is provably efficient in the context of general function approximation.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper34

问问它们各自怎么用它

它引用的顶会 Paper27

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖