A Clean Slate for Offline Reinforcement Learning
Matthew Thomas Jackson, Uljad Berdica, Jarek Liesen, Shimon Whiteson, Jakob N. Foerster
Abstract
Progress in offline reinforcement learning (RL) has been impeded by ambiguous problem definitions and entangled algorithmic designs, resulting in inconsistent implementations, insufficient ablations, and unfair evaluations. Although offline RL explicitly avoids environment interaction, prior methods frequently employ extensive, undocumented online evaluation for hyperparameter tuning, complicating method comparisons. Moreover, existing reference implementations differ significantly in boilerplate code, obscuring their core algorithmic contributions. We address these challenges by first introducing a rigorous taxonomy and a transparent evaluation protocol that explicitly quantifies online tuning budgets. To resolve opaque algorithmic design, we provide clean, minimalistic, singlefile implementations of various model-free and model-based offline RL methods, significantly enhancing clarity and achieving substantial speed-ups. Leveraging these streamlined implementations, we propose Unifloral, a unified algorithm that encapsulates diverse prior approaches within a single, comprehensive hyperparameter space, enabling algorithm development in a shared hyperparameter space. Using Unifloral with our rigorous evaluation protocol, we develop two novel algorithms-TD3-AWR (model-free) and MoBRAC (model-based)-which substantially outperform established baselines. Our implementation is publicly available at https://github.com/EmptyJackson/unifloral . * Equal contribution. Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 196d20ea-5c32-420a-b0ae-2b3d55044bd0Cited by top-tier papers1
Ask how each one uses itBuilds on19
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
Related papers
- Revisiting the Minimalist Approach to Offline Reinforcement LearningDenis Tarasov, Vladislav Kurenkov, Alexander Nikulin, Sergey KolesnikovNeurIPS 2023 · 148 citations
- Uni-RL: Unifying Online and Offline RL via Implicit Value RegularizationHaoran Xu, Liyuan Mao, Hui Jin, Weinan Zhang et al.NeurIPS 2025 · 3 citations
- Towards General-Purpose Model-Free Reinforcement LearningScott Fujimoto, Pierluca D'Oro, Amy Zhang, Yuandong Tian et al.ICLR 2025
- Reward-Consistent Dynamics Models are Strongly Generalizable for Offline Reinforcement LearningFan-Ming Luo, Tian Xu, Xingchen Cao, Yang YuICLR 2024 · 16 citations
- A Policy-Guided Imitation Approach for Offline Reinforcement LearningHaoran Xu, Li Jiang, Jianxiong Li, Xianyuan ZhanNeurIPS 2022 · 86 citations
