Provably Efficient Online Hyperparameter Optimization with Population-Based Bandits
Jack Parker-Holder, Vu Nguyen, Stephen J. Roberts
Abstract
Many of the recent triumphs in machine learning are dependent on well-tuned hyperparameters. This is particularly prominent in reinforcement learning (RL) where a small change in the configuration can lead to failure. Despite the importance of tuning hyperparameters, it remains expensive and is often done in a naive and laborious way. A recent solution to this problem is Population Based Training (PBT) which updates both weights and hyperparameters in a single training run of a population of agents. PBT has been shown to be particularly effective in RL, leading to widespread use in the field. However, PBT lacks theoretical guarantees since it relies on random heuristics to explore the hyperparameter space. This inefficiency means it typically requires vast computational resources, which is prohibitive for many small and medium sized labs. In this work, we introduce the first provably efficient PBT-style algorithm, Population-Based Bandits (PB2). PB2 uses a probabilistic model to guide the search in an efficient way, making it possible to discover high performing hyperparameter configurations with far fewer agents than typically required by PBT. We show in a series of RL experiments that PB2 is able to achieve high performance with a modest computational budget.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e123eeaf-a118-4bdd-b72f-c7acbbf81cb8Cited by top-tier papers20
- Hyperparameters in Reinforcement Learning and How To Tune ThemTheresa Eimer, Marius Lindauer, Roberta RaileanuICML 2023 · 96 citations
- Revisiting Design Choices in Offline Model Based Reinforcement LearningCong Lu, Philip J. Ball, Jack Parker-Holder, Michael A. Osborne et al.ICLR 2022 · 65 citations
- Sample-Efficient Automated Deep Reinforcement LearningJörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank HutterICLR 2021 · 49 citations
- Optimal Transport Kernels for Sequential and Parallel Neural Architecture SearchVu Nguyen, Tam Le, Makoto Yamada, Michael A. OsborneICML 2021 · 42 citations
- Bayesian Optimization for Iterative LearningVu Nguyen, Sebastian Schulze, Michael A. OsborneNeurIPS 2020 · 38 citations
Builds on3
- Bayesian Optimisation over Multiple Continuous and Categorical InputsBin Xin Ru, Ahsan S. Alvi, Vu Nguyen, Michael A. Osborne et al.ICML 2020 · 119 citations
- Knowing The What But Not The Where in Bayesian OptimizationVu Nguyen, Michael A. OsborneICML 2020 · 42 citations
- Bayesian Optimization for Iterative LearningVu Nguyen, Sebastian Schulze, Michael A. OsborneNeurIPS 2020 · 38 citations
Related papers
- Tuning Mixed Input Hyperparameters on the Fly for Efficient Population Based AutoRLJack Parker-Holder, Vu Nguyen, Shaan Desai, Stephen J. RobertsNeurIPS 2021 · 22 citations
- Iterated Population Based Training with Task-Agnostic RestartsAlexander Chebykin, Tanja Alderliesten, Peter A.N BosmanICML 2026
- Multi-Objective Population Based TrainingArkadiy Dushatskiy, Alexander Chebykin, Tanja Alderliesten, Peter A. N. BosmanICML 2023 · 4 citations
- Accelerating and Improving AlphaZero Using Population Based TrainingTi-Rong Wu, Ting-Han Wei, I-Chen WuAAAI 2020 · 19 citations
- Fast Population-Based Reinforcement Learning on a Single MachineArthur Flajolet, Claire Bizon Monroc, Karim Beguir, Thomas PierrotICML 2022 · 11 citations
