Transformers Can Do Bayesian Inference
Samuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka, Frank Hutter
摘要
Currently, it is hard to reap the benefits of deep learning for Bayesian methods, which allow the explicit specification of prior knowledge and accurately capture model uncertainty. We present Prior-Data Fitted Networks (PFNs). PFNs leverage in-context learning in large-scale machine learning techniques to approximate a large set of posteriors. The only requirement for PFNs to work is the ability to sample from a prior distribution over supervised learning tasks (or functions). Our method restates the objective of posterior approximation as a supervised classification problem with a set-valued input: it repeatedly draws a task (or function) from the prior, draws a set of data points and their labels from it, masks one of the labels and learns to make probabilistic predictions for it based on the set-valued input of the rest of the data points. Presented with a set of samples from a new supervised learning task as input, PFNs make probabilistic predictions for arbitrary other data points in a single forward propagation, having learned to approximate Bayesian inference. We demonstrate that PFNs can near-perfectly mimic Gaussian processes and also enable efficient Bayesian inference for intractable problems, with over 200-fold speedups in multiple setups compared to current methods. We obtain strong results in very diverse areas such as Gaussian process regression, Bayesian neural networks, classification for small tabular data sets, and few-shot image classification, demonstrating the generality of PFNs. Code and trained PFNs are released at https://github.com/ automl/TransformersCanDoBayesianInference .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper113
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence ModelingTung Nguyen, Aditya GroverICML 2022 · 被引用 148 次
- ForecastPFN: Synthetically-Trained Zero-Shot ForecastingSamuel Dooley, Gurnoor Singh Khurana, Chirag Mohapatra, Siddartha V. Naidu 等NeurIPS 2023 · 被引用 142 次
- TabDPT: Scaling Tabular Foundation Models on Real DataJunwei Ma, Valentin Thomas, Rasa Hosseinzadeh, Alex Labach 等NeurIPS 2025 · 被引用 118 次
- Towards Learning Universal Hyperparameter Optimizers with TransformersYutian Chen, Xingyou Song, Chansoo Lee, Zi Wang 等NeurIPS 2022 · 被引用 106 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- AugMix: A Simple Data Processing Method to Improve Robustness and UncertaintyDan Hendrycks, Norman Mu, Ekin Dogus Cubuk, Barret Zoph 等ICLR 2020 · 被引用 1,572 次
- BoTorch: A Framework for Efficient Monte-Carlo Bayesian OptimizationMaximilian Balandat, Brian Karrer, Daniel R. Jiang, Samuel Daulton 等NeurIPS 2020 · 被引用 686 次
相关 Paper
- Transformers Can Learn Posterior Predictive Distributions In-ContextGyeonghun Kang, Changwoo Lee, Xiang ChengICML 2026
- Statistical Foundations of Prior-Data Fitted NetworksThomas NaglerICML 2023 · 被引用 51 次
- PFNs4BO: In-Context Learning for Bayesian OptimizationSamuel Müller, Matthias Feurer, Noah Hollmann, Frank HutterICML 2023 · 被引用 71 次
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 被引用 96 次
- Efficient Bayesian Learning Curve Extrapolation using Prior-Data Fitted NetworksSteven Adriaensen, Herilalaina Rakotoarison, Samuel Müller, Frank HutterNeurIPS 2023 · 被引用 55 次
