Gray-Box Gaussian Processes for Automated Reinforcement Learning
Gresa Shala, André Biedenkapp, Frank Hutter, Josif Grabocka
Abstract
Despite having achieved spectacular milestones in an array of important realworld applications, most Reinforcement Learning (RL) methods are very brittle concerning their hyperparameters. Notwithstanding the crucial importance of setting the hyperparameters in training state-of-the-art agents, the task of hyperparameter optimization (HPO) in RL is understudied. In this paper, we propose a novel gray-box Bayesian Optimization technique for HPO in RL, that enriches Gaussian Processes with reward curve estimations based on generalized logistic functions. We thus not only reason about the performance of learning algorithms, transferring information across configurations but also about epochs of the learning algorithm. In a very large-scale experimental protocol, comprising 5 popular RL methods (DDPG, A2C, PPO, SAC, TD3), 22 environments (OpenAI Gym: Mujoco, Atari, Classic Control), and 7 HPO baselines, we demonstrate that our method significantly outperforms current HPO practices in RL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and HowSebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter et al.ICLR 2024 · 27 citations
- Efficient Cross-Episode Meta-RLGresa Shala, André Biedenkapp, Pierre Krack, Florian Walter et al.ICLR 2025
- Adaptive Q-Network: On-the-fly Target Selection for Deep Reinforcement LearningThéo Vincent, Fabian Wahren, Jan Peters, Boris Belousov et al.ICLR 2025
Builds on7
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- Provably Efficient Online Hyperparameter Optimization with Population-Based BanditsJack Parker-Holder, Vu Nguyen, Stephen J. RobertsNeurIPS 2020 · 105 citations
- Quantifying Differences in Reward FunctionsAdam Gleave, Michael Dennis, Shane Legg, Stuart Russell et al.ICLR 2021 · 77 citations
- What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale StudyMarcin Andrychowicz, Anton Raichuk, Piotr Stanczyk, Manu Orsini et al.ICLR 2021 · 52 citations
- Sample-Efficient Automated Deep Reinforcement LearningJörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank HutterICLR 2021 · 49 citations
Related papers
- Meta-Learning Acquisition Functions for Transfer Learning in Bayesian OptimizationMichael Volpp, Lukas P. Fröhlich, Kirsten Fischer, Andreas Doerr et al.ICLR 2020 · 104 citations
- Bayesian Optimization for Iterative LearningVu Nguyen, Sebastian Schulze, Michael A. OsborneNeurIPS 2020 · 38 citations
- Supervising the Multi-Fidelity Race of Hyperparameter ConfigurationsMartin Wistuba, Arlind Kadra, Josif GrabockaNeurIPS 2022 · 24 citations
- EARL-BO: Reinforcement Learning for Multi-Step Lookahead, High-Dimensional Bayesian OptimizationMujin Cheon, Jay H. Lee, Dong-Yeun Koh, Calvin TsayICML 2025
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 67 citations
