Gray-Box Gaussian Processes for Automated Reinforcement Learning
Gresa Shala, André Biedenkapp, Frank Hutter, Josif Grabocka
摘要
Despite having achieved spectacular milestones in an array of important realworld applications, most Reinforcement Learning (RL) methods are very brittle concerning their hyperparameters. Notwithstanding the crucial importance of setting the hyperparameters in training state-of-the-art agents, the task of hyperparameter optimization (HPO) in RL is understudied. In this paper, we propose a novel gray-box Bayesian Optimization technique for HPO in RL, that enriches Gaussian Processes with reward curve estimations based on generalized logistic functions. We thus not only reason about the performance of learning algorithms, transferring information across configurations but also about epochs of the learning algorithm. In a very large-scale experimental protocol, comprising 5 popular RL methods (DDPG, A2C, PPO, SAC, TD3), 22 environments (OpenAI Gym: Mujoco, Atari, Classic Control), and 7 HPO baselines, we demonstrate that our method significantly outperforms current HPO practices in RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and HowSebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter 等ICLR 2024 · 被引用 27 次
- Efficient Cross-Episode Meta-RLGresa Shala, André Biedenkapp, Pierre Krack, Florian Walter 等ICLR 2025
- Adaptive Q-Network: On-the-fly Target Selection for Deep Reinforcement LearningThéo Vincent, Fabian Wahren, Jan Peters, Boris Belousov 等ICLR 2025
它引用的顶会 Paper7
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 305 次
- Provably Efficient Online Hyperparameter Optimization with Population-Based BanditsJack Parker-Holder, Vu Nguyen, Stephen J. RobertsNeurIPS 2020 · 被引用 105 次
- Quantifying Differences in Reward FunctionsAdam Gleave, Michael Dennis, Shane Legg, Stuart Russell 等ICLR 2021 · 被引用 77 次
- What Matters for On-Policy Deep Actor-Critic Methods? A Large-Scale StudyMarcin Andrychowicz, Anton Raichuk, Piotr Stanczyk, Manu Orsini 等ICLR 2021 · 被引用 52 次
- Sample-Efficient Automated Deep Reinforcement LearningJörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank HutterICLR 2021 · 被引用 49 次
相关 Paper
- Meta-Learning Acquisition Functions for Transfer Learning in Bayesian OptimizationMichael Volpp, Lukas P. Fröhlich, Kirsten Fischer, Andreas Doerr 等ICLR 2020 · 被引用 104 次
- Bayesian Optimization for Iterative LearningVu Nguyen, Sebastian Schulze, Michael A. OsborneNeurIPS 2020 · 被引用 38 次
- Supervising the Multi-Fidelity Race of Hyperparameter ConfigurationsMartin Wistuba, Arlind Kadra, Josif GrabockaNeurIPS 2022 · 被引用 24 次
- EARL-BO: Reinforcement Learning for Multi-Step Lookahead, High-Dimensional Bayesian OptimizationMujin Cheon, Jay H. Lee, Dong-Yeun Koh, Calvin TsayICML 2025
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 被引用 67 次
