Few-Shot Bayesian Optimization with Deep Kernel Surrogates
Martin Wistuba, Josif Grabocka
Abstract
Hyperparameter optimization (HPO) is a central pillar in the automation of machine learning solutions and is mainly performed via Bayesian optimization, where a parametric surrogate is learned to approximate the black box response function (e.g. validation error). Unfortunately, evaluating the response function is computationally intensive. As a remedy, earlier work emphasizes the need for transfer learning surrogates which learn to optimize hyperparameters for an algorithm from other tasks. In contrast to previous work, we propose to rethink HPO as a few-shot learning problem in which we train a shared deep surrogate model to quickly adapt (with few response evaluations) to the response function of a new task. We propose the use of a deep kernel network for a Gaussian process surrogate that is meta-learned in an end-to-end fashion in order to jointly approximate the response functions of a collection of training data sets. As a result, the novel few-shot optimization of our deep kernel surrogate leads to new state-of-the-art results at HPO compared to several recent methods on diverse metadata sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3117b7db-6ebf-4779-bf38-0fb3a50d3ec4Cited by top-tier papers32
- Large Language Models to Enhance Bayesian OptimizationTennison Liu, Nicolás Astorga, Nabeel Seedat, Mihaela van der SchaarICLR 2024 · 143 citations
- Towards Learning Universal Hyperparameter Optimizers with TransformersYutian Chen, Xingyou Song, Chansoo Lee, Zi Wang et al.NeurIPS 2022 · 106 citations
- PFNs4BO: In-Context Learning for Bayesian OptimizationSamuel Müller, Matthias Feurer, Noah Hollmann, Frank HutterICML 2023 · 71 citations
- PriorBand: Practical Hyperparameter Optimization in the Age of Deep LearningNeeratyoy Mallik, Edward Bergman, Carl Hvarfner, Danny Stoll et al.NeurIPS 2023 · 50 citations
- End-to-End Meta-Bayesian Optimisation with Transformer Neural ProcessesAlexandre Maraval, Matthieu Zimmer, Antoine Grosnit, Haitham Bou-AmmarNeurIPS 2023 · 41 citations
Builds on3
- Bayesian Meta-Learning for the Few-Shot Setting via Deep KernelsMassimiliano Patacchiola, Jack Turner, Elliot J. Crowley, Michael F. P. O'Boyle et al.NeurIPS 2020 · 167 citations
- Meta-Learning Acquisition Functions for Transfer Learning in Bayesian OptimizationMichael Volpp, Lukas P. Fröhlich, Kirsten Fischer, Andreas Doerr et al.ICLR 2020 · 104 citations
- Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian ProcessesHao Chen, Lili Zheng, Raed Al Kontar, Garvesh RaskuttiNeurIPS 2020 · 49 citations
Related papers
- FSEO: Few-Shot Evolutionary Optimization via Meta-Learning for Expensive Multi-Objective OptimizationXunzhao YuNeurIPS 2025
- Deep Ranking Ensembles for Hyperparameter OptimizationAbdus Salam Khazi, Sebastian Pineda-Arango, Josif GrabockaICLR 2023 · 1 citation
- Informed Initialization for Bayesian Optimization and Active LearningCarl Hvarfner, David Eriksson, Eytan Bakshy, Maximilian BalandatNeurIPS 2025 · 4 citations
- Transfer NAS with Meta-learned Bayesian SurrogatesGresa Shala, Thomas Elsken, Frank Hutter, Josif GrabockaICLR 2023
- Meta-learning Hyperparameter Performance Prediction with Neural ProcessesYing Wei, Peilin Zhao, Junzhou HuangICML 2021 · 26 citations
