Batch Reinforcement Learning with Hyperparameter Gradients
Byung-Jun Lee, Jongmin Lee, Peter Vrancx, Dongho Kim, Kee-Eung Kim
Abstract
We consider the batch reinforcement learning problem where the agent needs to learn only from a fixed batch of data, without further interaction with the environment. In such a scenario, we want to prevent the optimized policy from deviating too much from the data collection policy since the estimation becomes highly unstable otherwise due to the off-policy nature of the problem. However, imposing this requirement too strongly will result in a policy that merely follows the data collection policy. Unlike prior work where this trade-off is controlled by hand-tuned hyperparameters, we propose a novel batch reinforcement learning approach, batch optimization of policy and hyperparameter (BOPAH), that uses a gradient-based optimization of the hyperparameter using held-out data. We show that BOPAH outperforms other batch reinforcement learning algorithms in tabular and continuous control tasks, by finding a good balance to the trade-off between adhering to the data collection policy and pursuing the possible policy improvement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Pessimistic Bootstrapping for Uncertainty-Driven Offline Reinforcement LearningChenjia Bai, Lingxiao Wang, Zhuoran Yang, Zhi-Hong Deng et al.ICLR 2022 · 173 citations
- OptiDICE: Offline Policy Optimization via Stationary Distribution Correction EstimationJongmin Lee, Wonseok Jeon, Byung-Jun Lee, Joelle Pineau et al.ICML 2021 · 137 citations
- COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction EstimationJongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz, Nicolas Heess et al.ICLR 2022 · 84 citations
- LobsDICE: Offline Learning from Observation via Stationary Distribution Correction EstimationGeon-Hyeong Kim, Jongmin Lee, Youngsoo Jang, Hongseok Yang et al.NeurIPS 2022 · 33 citations
- Bi-Level Offline Policy Optimization with Limited ExplorationWenzhuo ZhouNeurIPS 2023 · 6 citations
Builds on2
- An Optimistic Perspective on Offline Reinforcement LearningRishabh Agarwal, Dale Schuurmans, Mohammad NorouziICML 2020 · 568 citations
- Keep Doing What Worked: Behavior Modelling Priors for Offline Reinforcement LearningNoah Y. Siegel, Jost Tobias Springenberg, Felix Berkenkamp, Abbas Abdolmaleki et al.ICLR 2020 · 299 citations
Related papers
- Batch Reinforcement Learning Through Continuation MethodYijie Guo, Shengyu Feng, Nicolas Le Roux, Ed H. Chi et al.ICLR 2021 · 16 citations
- Continuous Doubly Constrained Batch Reinforcement LearningRasool Fakoor, Jonas Mueller, Kavosh Asadi, Pratik Chaudhari et al.NeurIPS 2021 · 37 citations
- Critic Regularized RegressionZiyu Wang, Alexander Novikov, Konrad Zolna, Josh Merel et al.NeurIPS 2020 · 406 citations
- Behavior Proximal Policy OptimizationZifeng Zhuang, Kun Lei, Jinxin Liu, Donglin Wang et al.ICLR 2023 · 8 citations
- BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement LearningXinyue Chen, Zijian Zhou, Zheng Wang, Che Wang et al.NeurIPS 2020 · 146 citations
