Tuning Mixed Input Hyperparameters on the Fly for Efficient Population Based AutoRL
Jack Parker-Holder, Vu Nguyen, Shaan Desai, Stephen J. Roberts
Abstract
Despite a series of recent successes in reinforcement learning (RL), many RL algorithms remain sensitive to hyperparameters. As such, there has recently been interest in the field of AutoRL, which seeks to automate design decisions to create more general algorithms. Recent work suggests that population based approaches may be effective AutoRL algorithms, by learning hyperparameter schedules on the fly. In particular, the PB2 algorithm is able to achieve strong performance in RL tasks by formulating online hyperparameter optimization as time varying GP-bandit problem, while also providing theoretical guarantees. However, PB2 is only designed to work for continuous hyperparameters, which severely limits its utility in practice. In this paper we introduce a new (provably) efficient hierarchical approach for optimizing both continuous and categorical variables, using a new time-varying bandit algorithm specifically designed for the population based training regime. We evaluate our approach on the challenging Procgen benchmark, where we show that explicitly modelling dependence between data augmentation and other hyperparameters improves generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7b847a7-265c-49b9-868f-7cf83c9e2080Cited by top-tier papers3
- Mixed-Variable Black-Box Optimisation Using Value Proposal TreesYan Zuo, Vu Nguyen, Amir Dezfouli, David Alexander et al.AAAI 2023
- Iterated Population Based Training with Task-Agnostic RestartsAlexander Chebykin, Tanja Alderliesten, Peter A.N BosmanICML 2026
- ULTHO: Ultra-Lightweight Yet Efficient Hyperparameter Optimization in Deep Reinforcement LearningMingqi Yuan, Bo Li, Xin Jin, Wenjun ZengICCV 2025
Builds on19
- CURL: Contrastive Unsupervised Representations for Reinforcement LearningMichael Laskin, Aravind Srinivas, Pieter AbbeelICML 2020 · 1,261 citations
- Image Augmentation Is All You Need: Regularizing Deep Reinforcement Learning from PixelsDenis Yarats, Ilya Kostrikov, Rob FergusICLR 2021 · 911 citations
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 685 citations
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras et al.ICLR 2020 · 305 citations
- AutoML-Zero: Evolving Machine Learning Algorithms From ScratchEsteban Real, Chen Liang, David R. So, Quoc V. LeICML 2020 · 265 citations
Related papers
- Provably Efficient Online Hyperparameter Optimization with Population-Based BanditsJack Parker-Holder, Vu Nguyen, Stephen J. RobertsNeurIPS 2020 · 105 citations
- Sample-Efficient Automated Deep Reinforcement LearningJörg K. H. Franke, Gregor Köhler, André Biedenkapp, Frank HutterICLR 2021 · 49 citations
- Automatic Data Augmentation for Generalization in Reinforcement LearningRoberta Raileanu, Maxwell Goldstein, Denis Yarats, Ilya Kostrikov et al.NeurIPS 2021 · 143 citations
- Multi-agent Dynamic Algorithm ConfigurationKe Xue, Jiacheng Xu, Lei Yuan, Miqing Li et al.NeurIPS 2022 · 65 citations
- On Effective Scheduling of Model-based Reinforcement LearningHang Lai, Jian Shen, Weinan Zhang, Yimin Huang et al.NeurIPS 2021 · 23 citations
