Optimal Uniform OPE and Model-based Offline Reinforcement Learning in Time-Homogeneous, Reward-Free and Task-Agnostic Settings
Ming Yin, Yu-Xiang Wang
摘要
This work studies the statistical limits of uniform convergence for offline policy evaluation (OPE) problems with model-based methods (for episodic MDP) and provides a unified framework towards optimal learning for several well-motivated offline tasks. Uniform OPE is a stronger measure than the point-wise OPE and ensures offline learning when contains all policies (the global class). In this paper, we establish an lower bound (over model-based family) for the global uniform OPE and our main result establishes an upper bound of for the local uniform convergence that applies to all near-empirically optimal policies for the MDPs with stationary transition. Here is the minimal marginal state-action probability. Critically, the highlight in achieving the optimal rate is our design of singleton absorbing MDP, which is a new sharp analysis tool that works with the model-based approach. We generalize such a model-based framework to the new settings: offline task-agnostic and the offline reward-free with optimal complexity ( is the number of tasks) and respectively. These results provide a unified solution for simultaneously solving different offline RL problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Towards Instance-Optimal Offline Reinforcement Learning with PessimismMing Yin, Yu-Xiang WangNeurIPS 2021 · 被引用 93 次
- Near-optimal Offline Reinforcement Learning with Linear Representation: Leveraging Variance Information with PessimismMing Yin, Yaqi Duan, Mengdi Wang, Yu-Xiang WangICLR 2022 · 被引用 74 次
- Offline Reinforcement Learning with Differential PrivacyDan Qiao, Yu-Xiang WangNeurIPS 2023 · 被引用 34 次
- Provable Benefit of Multitask Representation Learning in Reinforcement LearningYuan Cheng, Songtao Feng, Jing Yang, Hong Zhang 等NeurIPS 2022 · 被引用 33 次
- On the Statistical Efficiency of Reward-Free Exploration in Non-Linear RLJinglin Chen, Aditya Modi, Akshay Krishnamurthy, Nan Jiang 等NeurIPS 2022 · 被引用 31 次
它引用的顶会 Paper21
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- Is Pessimism Provably Efficient for Offline RL?Ying Jin, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 419 次
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang 等ICML 2020 · 被引用 324 次
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
相关 Paper
- Nearly Horizon-Free Offline Reinforcement LearningTongzheng Ren, Jialian Li, Bo Dai, Simon S. Du 等NeurIPS 2021 · 被引用 54 次
- Near-Optimal Offline Reinforcement Learning via Double Variance ReductionMing Yin, Yu Bai, Yu-Xiang WangNeurIPS 2021 · 被引用 72 次
- Asymptotically Exact Error Characterization of Offline Policy Evaluation with Misspecified Linear ModelsKohei MiyaguchiNeurIPS 2021 · 被引用 3 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- Optimal Single-Policy Sample Complexity and Transient Coverage for Average-Reward Offline RLMatthew Zurek, Guy Zamir, Yudong ChenNeurIPS 2025 · 被引用 2 次
