DeepAveragers: Offline Reinforcement Learning By Solving Derived Non-Parametric MDPs
Aayam Kumar Shrestha, Stefan Lee, Prasad Tadepalli, Alan Fern
摘要
We study an approach to offline reinforcement learning (RL) based on optimally solving finitely-represented MDPs derived from a static dataset of experience. This approach can be applied on top of any learned representation and has the potential to easily support multiple solution objectives as well as zero-shot adjustment to changing environments and goals. Our main contribution is to introduce the Deep Averagers with Costs MDP (DAC-MDP) and to investigate its solutions for offline RL. DAC-MDPs are a non-parametric model that can leverage deep representations and account for limited data by introducing costs for exploiting under-represented parts of the model. In theory, we show conditions that allow for lower-bounding the performance of DAC-MDP solutions. We also investigate the empirical behavior in a number of environments, including those with image-based observations. Overall, the experiments demonstrate that the framework can work in practice and scale to large complex offline RL problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Large-Scale Retrieval for Reinforcement LearningPeter Conway Humphreys, Arthur Guez, Olivier Tieleman, Laurent Sifre 等NeurIPS 2022 · 被引用 39 次
- Provably Efficient Offline Reinforcement Learning with Perturbed Data SourcesChengshuai Shi, Wei Xiong, Cong Shen, Jing YangICML 2023 · 被引用 5 次
- Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement LearningGuy Azran, Mohamad H. Danesh, Stefano V. Albrecht, Sarah KerenAAAI 2024 · 被引用 2 次
- Value Memory Graph: A Graph-Structured World Model for Offline Reinforcement LearningDeyao Zhu, Li Erran Li, Mohamed ElhoseinyICLR 2023
- Efficient Exploration and Discriminative World Model Learning with an Object-Centric AbstractionAnthony GX-Chen, Kenneth Marino, Rob FergusICLR 2025
它引用的顶会 Paper8
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 被引用 2,881 次
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon 等NeurIPS 2020 · 被引用 989 次
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 被引用 870 次
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 被引用 568 次
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel 等NeurIPS 2020 · 被引用 406 次
相关 Paper
- Representation Balancing Offline Model-based Reinforcement LearningByung-Jun Lee, Jongmin Lee, Kee-Eung KimICLR 2021 · 被引用 8 次
- Revisiting the Linear-Programming Framework for Offline RL with General Function ApproximationAsuman E. Ozdaglar, Sarath Pattathil, Jiawei Zhang, Kaiqing ZhangICML 2023 · 被引用 8 次
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 被引用 47 次
- Offline Reinforcement Learning as Anti-explorationShideh Rezaeifar, Robert Dadashi, Nino Vieillard, Léonard Hussenot 等AAAI 2022 · 被引用 64 次
- Offline Meta-Reinforcement Learning with Advantage WeightingEric Mitchell, Rafael Rafailov, Xue Bin Peng, Sergey Levine 等ICML 2021 · 被引用 122 次
