Provably Good Batch Off-Policy Reinforcement Learning Without Great Exploration
Yao Liu, Adith Swaminathan, Alekh Agarwal, Emma Brunskill
摘要
Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is challenging: a new decision policy may visit states and actions outside the support of the batch data, and function approximation and optimization with limited samples can further increase the potential of learning policies with overly optimistic estimates of their future performance. Some recent approaches to address these concerns have shown promise, but can still be overly optimistic in their expected outcomes. Theoretical work that provides strong guarantees on the performance of the output policy relies on a strong concentrability assumption, which makes it unsuitable for cases where the ratio between state-action distributions of behavior policy and some candidate policies is large. This is because, in the traditional analysis, the error bound scales up with this ratio. We show that using pessimistic value estimates in the low-data regions in Bellman optimality and evaluation back-up can yield more adaptive and stronger guarantees when the concentrability assumption does not hold. In certain settings, they can find the approximately best policy within the state-action space explored by the batch data, without requiring a priori assumptions of concentrability. We highlight the necessity of our pessimistic update and the limitations of previous algorithms and analyses by illustrative MDP examples and demonstrate an empirical comparison of our algorithm and other state-of-the-art batch RL baselines in standard benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper49
- Bellman-consistent Pessimism for Offline Reinforcement LearningTengyang Xie, Ching-An Cheng, Nan Jiang, Paul Mineiro 等NeurIPS 2021 · 被引用 339 次
- Pessimistic Model-based Offline Reinforcement Learning under Partial CoverageMasatoshi Uehara, Wen SunICLR 2022 · 被引用 176 次
- Supervised Pretraining Can Learn In-Context Reinforcement LearningJonathan Lee, Annie Xie, Aldo Pacchiano, Yash Chandak 等NeurIPS 2023 · 被引用 170 次
- Representation Learning for Online and Offline RL in Low-rank MDPsMasatoshi Uehara, Xuezhou Zhang, Wen SunICLR 2022 · 被引用 138 次
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 被引用 130 次
它引用的顶会 Paper3
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- Batch Stationary Distribution EstimationJunfeng Wen, Bo Dai, Lihong Li, Dale SchuurmansICML 2020 · 被引用 25 次
相关 Paper
- Continuous Doubly Constrained Batch Reinforcement LearningRasool Fakoor, Jonas Mueller, Kavosh Asadi, Pratik Chaudhari 等NeurIPS 2021 · 被引用 37 次
- Pessimistic Q-Learning for Offline Reinforcement Learning: Towards Optimal Sample ComplexityLaixi Shi, Gen Li, Yuting Wei, Yuxin Chen 等ICML 2022 · 被引用 110 次
- Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of PessimismParia Rashidinejad, Banghua Zhu, Cong Ma, Jiantao Jiao 等NeurIPS 2021 · 被引用 373 次
- Information-Directed Pessimism for Offline Reinforcement LearningAlec Koppel, Sujay Bhatt, Jiacheng Guo, Joe Eappen 等ICML 2024 · 被引用 3 次
- On Instance-Dependent Bounds for Offline Reinforcement Learning with Linear Function ApproximationThanh Nguyen-Tang, Ming Yin, Sunil Gupta, Svetha Venkatesh 等AAAI 2023 · 被引用 24 次
