PFNs4BO: In-Context Learning for Bayesian Optimization
Samuel Müller, Matthias Feurer, Noah Hollmann, Frank Hutter
Abstract
In this paper, we use Prior-data Fitted Networks (PFNs) as a flexible surrogate for Bayesian Optimization (BO). PFNs are neural processes that are trained to approximate the posterior predictive distribution (PPD) through in-context learning on any prior distribution that can be efficiently sampled from. We describe how this flexibility can be exploited for surrogate modeling in BO. We use PFNs to mimic a naive Gaussian process (GP), an advanced GP, and a Bayesian Neural Network (BNN). In addition, we show how to incorporate further information into the prior, such as allowing hints about the position of optima (user priors), ignoring irrelevant dimensions, and performing non-myopic BO by learning the acquisition function. The flexibility underlying these extensions opens up vast possibilities for using PFNs for BO. We demonstrate the usefulness of PFNs for BO in a large-scale evaluation on artificial GP samples and three different hyperparameter optimization testbeds: HPO-B, Bayesmark, and PD1. We publish code alongside trained models at github.com/automl/PFNs4BO.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 701c748f-db9e-41b4-ae4f-8ca1c9d1f59eCited by top-tier papers30
- Large Language Models to Enhance Bayesian OptimizationTennison Liu, Nicolás Astorga, Nabeel Seedat, Mihaela van der SchaarICLR 2024 · 143 citations
- ForecastPFN: Synthetically-Trained Zero-Shot ForecastingSamuel Dooley, Gurnoor Singh Khurana, Chirag Mohapatra, Siddartha V. Naidu et al.NeurIPS 2023 · 142 citations
- A Study of Bayesian Neural Network Surrogates for Bayesian OptimizationYucen Lily Li, Tim G. J. Rudner, Andrew Gordon WilsonICLR 2024 · 59 citations
- Do-PFN: In-Context Learning for Causal Effect EstimationJake Robertson, Arik Reuter, Siyuan Guo, Noah Hollmann et al.NeurIPS 2025 · 58 citations
- TuneTables: Context Optimization for Scalable Prior-Data Fitted NetworksBenjamin Feuer, Robin Schirrmeister, Valeriia Cherepanova, Chinmay Hegde et al.NeurIPS 2024 · 57 citations
Builds on6
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka et al.ICLR 2022 · 287 citations
- Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence ModelingTung Nguyen, Aditya GroverICML 2022 · 148 citations
- Interpretable Neural Architecture Search via Bayesian Optimisation with Weisfeiler-Lehman KernelsBin Xin Ru, Xingchen Wan, Xiaowen Dong, Michael A. OsborneICLR 2021 · 116 citations
- Towards Learning Universal Hyperparameter Optimizers with TransformersYutian Chen, Xingyou Song, Chansoo Lee, Zi Wang et al.NeurIPS 2022 · 106 citations
- Few-Shot Bayesian Optimization with Deep Kernel SurrogatesMartin Wistuba, Josif GrabockaICLR 2021 · 87 citations
Related papers
- In-Context Freeze-Thaw Bayesian Optimization for Hyperparameter OptimizationHerilalaina Rakotoarison, Steven Adriaensen, Neeratyoy Mallik, Samir Garibov et al.ICML 2024 · 28 citations
- Bayesian Optimization via Continual Variational Last Layer TrainingPaul Brunzema, Mikkel Jordahn, John Willes, Sebastian Trimpe et al.ICLR 2025
- Objective Bound Conditional Gaussian Process for Bayesian OptimizationTaewon Jeong, Heeyoung KimICML 2021 · 3 citations
- Transformers Can Learn Posterior Predictive Distributions In-ContextGyeonghun Kang, Changwoo Lee, Xiang ChengICML 2026
- GIT-BO: High-Dimensional Bayesian Optimization with Tabular Foundation ModelsRosen Ting-Ying Yu, Cyril Picard, Faez AhmedICLR 2026 · 13 citations
