DeepAveragers: Offline Reinforcement Learning By Solving Derived Non-Parametric MDPs
Aayam Kumar Shrestha, Stefan Lee, Prasad Tadepalli, Alan Fern
Abstract
We study an approach to offline reinforcement learning (RL) based on optimally solving finitely-represented MDPs derived from a static dataset of experience. This approach can be applied on top of any learned representation and has the potential to easily support multiple solution objectives as well as zero-shot adjustment to changing environments and goals. Our main contribution is to introduce the Deep Averagers with Costs MDP (DAC-MDP) and to investigate its solutions for offline RL. DAC-MDPs are a non-parametric model that can leverage deep representations and account for limited data by introducing costs for exploiting under-represented parts of the model. In theory, we show conditions that allow for lower-bounding the performance of DAC-MDP solutions. We also investigate the empirical behavior in a number of environments, including those with image-based observations. Overall, the experiments demonstrate that the framework can work in practice and scale to large complex offline RL problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Large-Scale Retrieval for Reinforcement LearningPeter Conway Humphreys, Arthur Guez, Olivier Tieleman, Laurent Sifre et al.NeurIPS 2022 · 39 citations
- Provably Efficient Offline Reinforcement Learning with Perturbed Data SourcesChengshuai Shi, Wei Xiong, Cong Shen, Jing YangICML 2023 · 5 citations
- Contextual Pre-planning on Reward Machine Abstractions for Enhanced Transfer in Deep Reinforcement LearningGuy Azran, Mohamad H. Danesh, Stefano V. Albrecht, Sarah KerenAAAI 2024 · 2 citations
- Value Memory Graph: A Graph-Structured World Model for Offline Reinforcement LearningDeyao Zhu, Li Erran Li, Mohamed ElhoseinyICLR 2023
- Efficient Exploration and Discriminative World Model Learning with an Object-Centric AbstractionAnthony GX-Chen, Kenneth Marino, Rob FergusICLR 2025
Builds on8
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel et al.NeurIPS 2020 · 406 citations
Related papers
- Representation Balancing Offline Model-based Reinforcement LearningByung-Jun Lee, Jongmin Lee, Kee-Eung KimICLR 2021 · 8 citations
- Revisiting the Linear-Programming Framework for Offline RL with General Function ApproximationAsuman E. Ozdaglar, Sarath Pattathil, Jiawei Zhang, Kaiqing ZhangICML 2023 · 8 citations
- DARA: Dynamics-Aware Reward Augmentation in Offline Reinforcement LearningJinxin Liu, Hongyin Zhang, Donglin WangICLR 2022 · 47 citations
- Offline Reinforcement Learning as Anti-explorationShideh Rezaeifar, Robert Dadashi, Nino Vieillard, Léonard Hussenot et al.AAAI 2022 · 64 citations
- Offline Meta-Reinforcement Learning with Advantage WeightingEric Mitchell, Rafael Rafailov, Xue Bin Peng, Sergey Levine et al.ICML 2021 · 122 citations
