Efficient Autoregressive Inference for Transformer Probabilistic Models
Conor Hassan, Nasrulloh R. B. S. Loka, Cen-You Li, Daolang Huang, Paul E. Chang, Yang Yang, Francesco Silvestrin, Samuel Kaski, Luigi Acerbi
Abstract
Set-based transformer models for amortized probabilistic inference and meta-learning, such as neural processes, prior-fitted networks, and tabular foundation models, excel at single-pass marginal prediction. However, many applications require joint distributions over multiple predictions. Purely autoregressive architectures generate these efficiently but sacrifice flexible set-conditioning. Obtaining joint distributions from set-based models requires re-encoding the entire context at each autoregressive step, which scales poorly. We introduce a causal autoregressive buffer that combines the strengths of both paradigms. The model encodes the context once and caches it; a lightweight causal buffer captures dependencies among generated targets, with each new prediction attending to both the cached context and all previously predicted targets added to the buffer. This enables efficient batched autoregressive sampling and joint predictive density evaluation. Training integrates set-based and autoregressive modes through masked attention at minimal overhead. Across synthetic functions, EEG time series, a Bayesian model comparison task, and tabular regression, our method closely matches the performance of full context re-encoding while delivering up to faster joint sampling and density evaluation, and up to lower memory usage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d701460-ab99-4eea-a881-b3e7d282fcd9Cited by top-tier papers5
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 85 citations
- ALINE: Joint Amortization for Bayesian Inference and Active Data AcquisitionDaolang Huang, Xinyi Wen, Ayush Bharti, Samuel Kaski et al.NeurIPS 2025 · 8 citations
- In-Context Multi-Objective OptimizationXinyu Zhang, Conor Hassan, Julien Martinelli, Daolang Huang et al.ICLR 2026 · 6 citations
- Incremental Transformer Neural ProcessesPhilip Mortimer, Cristiana Diaconu, Tommy Rochussen, Bruno Mlodozeniec et al.ICML 2026
- Revisiting Neural Processes via Fourier Transform and Volterra SeriesPeiman Mohseni, Nick Duffield, Raymond K WongICML 2026
Builds on32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
Related papers
- Autoregressive Conditional Neural ProcessesWessel P. Bruinsma, Stratis Markou, James Requeima, Andrew Y. K. Foong et al.ICLR 2023 · 3 citations
- Multi-Task Bayesian In-Context LearningQingyang Zhu, Eric Oermann, Kyunghyun ChoICML 2026
- Global Perception Based Autoregressive Neural ProcessesJinyang TaiICCV 2023 · 1 citation
- Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence ModelingTung Nguyen, Aditya GroverICML 2022 · 148 citations
- Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior AdaptationGeorge Whittle, Juliusz Ziomek, Jacob Rawling, Michael A OsborneICML 2026 · 13 citations
