Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior Adaptation
George Whittle, Juliusz Ziomek, Jacob Rawling, Michael A Osborne
摘要
While Bayesian inference provides a principled framework for reasoning under uncertainty, its widespread adoption is limited by the intractability of exact posterior computation, necessitating the use of approximate inference. However, existing methods are often computationally expensive, or demand costly retraining when priors change, limiting their utility, particularly in sequential inference problems such as real-time sensor fusion. To address these challenges, we introduce the Distribution Transformer---a novel architecture that can learn arbitrary distribution-to-distribution mappings. Our method can be trained to map a prior to the corresponding posterior, conditioned on some dataset---thus performing approximate Bayesian inference. Our novel architecture represents a prior distribution as a (universally-approximating) Gaussian Mixture Model (GMM), and transforms it into a GMM representation of the posterior. The components of the GMM attend to each other via self-attention, and to the datapoints via cross-attention. We demonstrate that Distribution Transformers both maintain flexibility to vary the prior, and significantly reduces computation times—from minutes to milliseconds—while achieving expected log-likelihood performance on par with or superior to existing approximate inference methods across tasks such as sequential inference, quantum system parameter inference, and Gaussian Process predictive posterior inference with hyperpriors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Effortless, Simulation-Efficient Bayesian Inference using Tabular Foundation ModelsJulius Vetter, Manuel Glöckler, Daniel Gedon, Jakob H. MackeNeurIPS 2025 · 被引用 14 次
- ALINE: Joint Amortization for Bayesian Inference and Active Data AcquisitionDaolang Huang, Xinyi Wen, Ayush Bharti, Samuel Kaski 等NeurIPS 2025 · 被引用 8 次
- Efficient Autoregressive Inference for Transformer Probabilistic ModelsConor Hassan, Nasrulloh R. B. S. Loka, Cen-You Li, Daolang Huang 等ICLR 2026 · 被引用 5 次
- PriorGuide: Test-Time Prior Adaptation for Simulation-Based InferenceYang Yang, Severi Rissanen, Paul Edmund Chang, Nasrulloh Ratu Bagus Satrio Loka 等ICLR 2026 · 被引用 3 次
- -PFN: Fast Entropy Search via In-Context LearningHerilalaina Rakotoarison, Steven Adriaensen, Tom Viering, Carl Hvarfner 等ICML 2026
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- Flow Matching for Scalable Simulation-Based InferenceJonas Wildberger, Maximilian Dax, Simon Buchholz, Stephen R. Green 等NeurIPS 2023 · 被引用 153 次
- Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence ModelingTung Nguyen, Aditya GroverICML 2022 · 被引用 148 次
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 被引用 96 次
相关 Paper
- Multi-Task Bayesian In-Context LearningQingyang Zhu, Eric Oermann, Kyunghyun ChoICML 2026
- Calibrating Transformers via Sparse Gaussian ProcessesWenlong Chen, Yingzhen LiICLR 2023
- Transformers can optimally learn regression mixture modelsReese Pathak, Rajat Sen, Weihao Kong, Abhimanyu DasICLR 2024 · 被引用 16 次
- Transformers as Unsupervised Learning Algorithms: A study on Gaussian MixturesZhiheng Chen, Ruofan Wu, Guanhua FangICLR 2026 · 被引用 2 次
- TACTiS: Transformer-Attentional Copulas for Time SeriesAlexandre Drouin, Étienne Marcotte, Nicolas ChapadosICML 2022 · 被引用 55 次
