Can Transformers Learn Full Bayesian Inference in Context?
Arik Reuter, Tim G. J. Rudner, Vincent Fortuin, David Rügamer
Abstract
Transformers have emerged as the dominant architecture in the field of deep learning, with a broad range of applications and remarkable in-context learning (ICL) capabilities. While not yet fully understood, ICL has already proved to be an intriguing phenomenon, allowing transformers to learn in context-without requiring further training. In this paper, we further advance the understanding of ICL by demonstrating that transformers can perform full Bayesian inference for commonly used statistical models in context. More specifically, we introduce a general framework that builds on ideas from prior fitted networks and continuous normalizing flows and enables us to infer complex posterior distributions for models such as generalized linear models and latent factor models. Extensive experiments on real-world datasets demonstrate that our ICL approach yields posterior samples that are similar in quality to state-of-the-art MCMC or variational inference methods that do not operate in context.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 33f9b395-3cfd-41de-a951-b05d18de8e3dCited by top-tier papers8
- GIT-BO: High-Dimensional Bayesian Optimization with Tabular Foundation ModelsRosen Ting-Ying Yu, Cyril Picard, Faez AhmedICLR 2026 · 13 citations
- In-Context Learning Is Provably Bayesian Inference: A Generalization Theory for Meta-LearningTomoya Wakayama, Taiji SuzukiICML 2026 · 12 citations
- Predictable Compression Failures: Order Sensitivity and Information Budgeting for Evidence-Grounded Binary AdjudicationLeon Chlon, Ahmed Karim, MarcAntonio AwadaICML 2026 · 2 citations
- ANCHOR: Abductive Network Construction with Hierarchical Orchestration for Reliable Probability Inference in Large Language ModelsWentao Qiu, Guanran Luo, Zhongquan Jian, Jingqi Gao et al.ICML 2026 · 2 citations
- Attention Implements the Fisher Geometry of Exponential FamiliesBodie RubacherICML 2026
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
Related papers
- Multi-Task Bayesian In-Context LearningQingyang Zhu, Eric Oermann, Kyunghyun ChoICML 2026
- In-Context Learning through the Bayesian PrismMadhur Panwar, Kabir Ahuja, Navin GoyalICLR 2024 · 79 citations
- Transformers Can Learn Posterior Predictive Distributions In-ContextGyeonghun Kang, Changwoo Lee, Xiang ChengICML 2026
- When can in-context learning generalize out of task distribution?Page C. Goddard, Lindsay M. Smith, Vudtiwat Ngampruetikorn, David J. SchwabICML 2025
- How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?Jingfeng Wu, Difan Zou, Zixiang Chen, Vladimir Braverman et al.ICLR 2024 · 94 citations
