Towards Learning Universal Hyperparameter Optimizers with Transformers
Yutian Chen, Xingyou Song, Chansoo Lee, Zi Wang, Richard Zhang, David Dohan, Kazuya Kawakami, Greg Kochanski, Arnaud Doucet, Marc'Aurelio Ranzato, Sagi Perel, Nando de Freitas
Abstract
Meta-learning hyperparameter optimization (HPO) algorithms from prior experiments is a promising approach to improve optimization efficiency over objective functions from a similar distribution. However, existing methods are restricted to learning from experiments sharing the same set of hyperparameters. In this paper, we introduce the OptFormer, the first text-based Transformer HPO framework that provides a universal end-to-end interface for jointly learning policy and function prediction when trained on vast tuning data from the wild, such as Google's Vizier database, one of the world's largest HPO datasets. Our extensive experiments demonstrate that the OptFormer can simultaneously imitate at least 7 different HPO algorithms, which can be further improved via its function uncertainty estimates. Compared to a Gaussian Process, the OptFormer also learns a robust prior distribution for hyperparameter response functions, and can thereby provide more accurate and better calibrated predictions. This work paves the path to future extensions for training a Transformer-based model as a general HPO optimizer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50187967-e635-4c88-bb25-7901c800591cCited by top-tier papers26
- Large Language Models as OptimizersChengrun Yang, Xuezhi Wang, Yifeng Lu, Hanxiao Liu et al.ICLR 2024 · 817 citations
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand et al.ICML 2023 · 155 citations
- Large Language Models to Enhance Bayesian OptimizationTennison Liu, Nicolás Astorga, Nabeel Seedat, Mihaela van der SchaarICLR 2024 · 143 citations
- PFNs4BO: In-Context Learning for Bayesian OptimizationSamuel Müller, Matthias Feurer, Noah Hollmann, Frank HutterICML 2023 · 71 citations
- End-to-End Meta-Bayesian Optimisation with Transformer Neural ProcessesAlexandre Maraval, Matthieu Zimmer, Antoine Grosnit, Haitham Bou-AmmarNeurIPS 2023 · 41 citations
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- NAS-Bench-201: Extending the Scope of Reproducible Neural Architecture SearchXuanyi Dong, Yi YangICLR 2020 · 825 citations
Related papers
- Deep Ranking Ensembles for Hyperparameter OptimizationAbdus Salam Khazi, Sebastian Pineda-Arango, Josif GrabockaICLR 2023 · 1 citation
- Optimizing Hyperparameters with Conformal Quantile RegressionDavid Salinas, Jacek Golebiowski, Aaron Klein, Matthias W. Seeger et al.ICML 2023 · 13 citations
- BOFormer: Learning to Solve Multi-Objective Bayesian Optimization via Non-Markovian RLYu-Heng Hung, Kai-Jie Lin, Yu-Heng Lin, Chien-Yi Wang et al.ICLR 2025
- In-Context Learning through the Bayesian PrismMadhur Panwar, Kabir Ahuja, Navin GoyalICLR 2024 · 79 citations
- BO: Augmenting Acquisition Functions with User Beliefs for Bayesian OptimizationCarl Hvarfner, Danny Stoll, Artur L. F. Souza, Marius Lindauer et al.ICLR 2022 · 93 citations
