Sample-Efficient Automated Deep Reinforcement Learning
Jörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank Hutter
Abstract
Despite significant progress in challenging problems across various domains, applying state-of-the-art deep reinforcement learning (RL) algorithms remains challenging due to their sensitivity to the choice of hyperparameters. This sensitivity can partly be attributed to the non-stationarity of the RL problem, potentially requiring different hyperparameter settings at various stages of the learning process. Additionally, in the RL setting, hyperparameter optimization (HPO) requires a large number of environment interactions, hindering the transfer of the successes in RL to real-world applications. In this work, we tackle the issues of sample-efficient and dynamic HPO in RL. We propose a population-based automated RL (AutoRL) framework to meta-optimize arbitrary off-policy RL algorithms. In this framework, we optimize the hyperparameters and also the neural architecture while simultaneously training the agent. By sharing the collected experience across the population, we substantially increase the sample efficiency of the meta-optimization. We demonstrate the capabilities of our sample-efficient AutoRL approach in a case study with the popular TD3 algorithm in the MuJoCo benchmark suite, where we reduce the number of environment interactions needed for meta-optimization by up to an order of magnitude compared to population-based training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8430dd3-cf20-4cfc-a97f-c8fbf7f41c6bCited by top-tier papers13
- Hyperparameters in Reinforcement Learning and How To Tune ThemTheresa Eimer, Marius Lindauer, Roberta RaileanuICML 2023 · 96 citations
- A Method for Evaluating Hyperparameter Sensitivity in Reinforcement LearningJacob Adkins, Michael Bowling, Adam WhiteNeurIPS 2024 · 36 citations
- Reinforcement Learning with Automated Auxiliary Loss SearchTairan He, Yuge Zhang, Kan Ren, Minghuan Liu et al.NeurIPS 2022 · 22 citations
- Tuning Mixed Input Hyperparameters on the Fly for Efficient Population Based AutoRLJack Parker-Holder, Vu Nguyen, Shaan Desai, Stephen J. RobertsNeurIPS 2021 · 22 citations
- Evolving Reinforcement Learning AlgorithmsJohn D. Co-Reyes, Yingjie Miao, Daiyi Peng, Esteban Real et al.ICLR 2021 · 19 citations
Builds on4
- A Self-Tuning Actor-Critic AlgorithmTom Zahavy, Zhongwen Xu, Vivek Veeriah, Matteo Hessel et al.NeurIPS 2020 · 106 citations
- Provably Efficient Online Hyperparameter Optimization with Population-Based BanditsJack Parker-Holder, Vu Nguyen, Stephen J. RobertsNeurIPS 2020 · 105 citations
- Proximal Distilled Evolutionary Reinforcement LearningCristian Bodnar, Ben Day, Pietro LióAAAI 2020 · 101 citations
- Off-Policy Actor-Critic with Shared Experience ReplaySimon Schmitt, Matteo Hessel, Karen SimonyanICML 2020 · 71 citations
Related papers
- Adaptive Q-Network: On-the-fly Target Selection for Deep Reinforcement LearningThéo Vincent, Fabian Wahren, Jan Peters, Boris Belousov et al.ICLR 2025
- Gray-Box Gaussian Processes for Automated Reinforcement LearningGresa Shala, André Biedenkapp, Frank Hutter, Josif GrabockaICLR 2023
- Meta-Gradient Reinforcement Learning with an Objective Discovered OnlineZhongwen Xu, Hado Philip van Hasselt, Matteo Hessel, Junhyuk Oh et al.NeurIPS 2020 · 90 citations
- ULTHO: Ultra-Lightweight Yet Efficient Hyperparameter Optimization in Deep Reinforcement LearningMingqi Yuan, Bo Li, Xin Jin, Wenjun ZengICCV 2025
- Iterated Population Based Training with Task-Agnostic RestartsAlexander Chebykin, Tanja Alderliesten, Peter A.N BosmanICML 2026
