The Difficulty of Passive Learning in Deep Reinforcement Learning
Georg Ostrovski, Pablo Samuel Castro, Will Dabney
摘要
Learning to act from observational data without active environmental interaction is a well-known challenge in Reinforcement Learning (RL). Recent approaches involve constraints on the learned policy or conservative updates, preventing strong deviations from the state-action distribution of the dataset. Although these methods are evaluated using non-linear function approximation, theoretical justifications are mostly limited to the tabular or linear cases. Given the impressive results of deep reinforcement learning, we argue for a need to more clearly understand the challenges in this setting. In the vein of Held & Hein's classic 1963 experiment, we propose the "tandem learning" experimental paradigm which facilitates our empirical analysis of the difficulties in offline reinforcement learning. We identify function approximation in conjunction with fixed data distributions as the strongest factors, thereby extending but also challenging hypotheses stated in past work. Our results provide relevant insights for offline deep reinforcement learning, while also shedding new light on phenomena observed in the online case of learning control.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper29
- Bigger, Better, Faster: Human-level Atari with human-level efficiencyMax Schwarzer, Johan S. Obando-Ceron, Aaron C. Courville, Marc G. Bellemare 等ICML 2023 · 被引用 155 次
- Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate ProgressRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2022 · 被引用 95 次
- Mixtures of Experts Unlock Parameter Scaling for Deep RLJohan S. Obando-Ceron, Ghada Sokar, Timon Willi, Clare Lyle 等ICML 2024 · 被引用 74 次
- Cross-Domain Policy Adaptation via Value-Guided Data FilteringKang Xu, Chenjia Bai, Xiaoteng Ma, Dong Wang 等NeurIPS 2023 · 被引用 41 次
- Learning Dynamics and Generalization in Deep Reinforcement LearningClare Lyle, Mark Rowland, Will Dabney, Marta Kwiatkowska 等ICML 2022 · 被引用 40 次
它引用的顶会 Paper15
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- Uncertainty Weighted Actor-Critic for Offline Reinforcement LearningYue Wu, Shuangfei Zhai, Nitish Srivastava, Joshua M. Susskind 等ICML 2021 · 被引用 223 次
相关 Paper
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 被引用 127 次
- Regularizing a Model-based Policy Stationary Distribution to Stabilize Offline Reinforcement LearningShentao Yang, Yihao Feng, Shujian Zhang, Mingyuan ZhouICML 2022 · 被引用 14 次
- Instabilities of Offline RL with Pre-Trained Neural RepresentationRuosong Wang, Yifan Wu, Ruslan Salakhutdinov, Sham M. KakadeICML 2021 · 被引用 46 次
- What are the Statistical Limits of Offline RL with Linear Function Approximation?Ruosong Wang, Dean P. Foster, Sham M. KakadeICLR 2021 · 被引用 172 次
