Lune

AAAI2024Top-tier venue

Exact Policy Recovery in Offline RL with Both Heavy-Tailed Rewards and Data Corruption

Yiding Chen, Xuezhou Zhang, Qiaomin Xie, Xiaojin Zhu

2024Year
2Citations
2Top-tier citations

Abstract

We study offline reinforcement learning (RL) with heavytailed reward distribution and data corruption: (i) Moving beyond subGaussian reward distribution, we require only a bounded (1 + γ)-th moment for γ ∈ (0, 1]; (ii) We allow corruptions where an attacker can arbitrarily modify -fraction of the rewards and transitions in the dataset. We first derive a sufficient optimality condition for generalized Pessimistic Value Iteration (PEVI), which allows various estimators with proper confidence bounds and can be applied to multiple learning settings. In order to handle the data corruption and heavy-tailed reward setting, we prove that the trimmed-mean estimation achieves a minimax optimal error rate O σ γ 1+γ

for robust mean estimation under heavy-tailed distributions. In the PEVI algorithm, we plug in the trimmed mean estimation and the confidence bound to solve the robust offline RL problem. Standard analysis reveals that data corruption induces a bias term O Hσ γ 1+γ + H in the suboptimality gap, which gives the false impression that any data corruption prevents optimal policy learning. By using the optimality condition for the generalized PEVI, we show that as long as the bias term is less than the "action gap", the policy returned by PEVI achieves the optimal value given sufficient data.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e83d03b4-2a22-4e0e-84a0-7697facff19b

Cited by top-tier papers2

Ask how each one uses it

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines