Efficient Autoregressive Inference for Transformer Probabilistic Models
Conor Hassan, Nasrulloh R. B. S. Loka, Cen-You Li, Daolang Huang, Paul E. Chang, Yang Yang, Francesco Silvestrin, Samuel Kaski, Luigi Acerbi
摘要
Set-based transformer models for amortized probabilistic inference and meta-learning, such as neural processes, prior-fitted networks, and tabular foundation models, excel at single-pass marginal prediction. However, many applications require joint distributions over multiple predictions. Purely autoregressive architectures generate these efficiently but sacrifice flexible set-conditioning. Obtaining joint distributions from set-based models requires re-encoding the entire context at each autoregressive step, which scales poorly. We introduce a causal autoregressive buffer that combines the strengths of both paradigms. The model encodes the context once and caches it; a lightweight causal buffer captures dependencies among generated targets, with each new prediction attending to both the cached context and all previously predicted targets added to the buffer. This enables efficient batched autoregressive sampling and joint predictive density evaluation. Training integrates set-based and autoregressive modes through masked attention at minimal overhead. Across synthetic functions, EEG time series, a Bayesian model comparison task, and tabular regression, our method closely matches the performance of full context re-encoding while delivering up to faster joint sampling and density evaluation, and up to lower memory usage.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- TabICLv2: A Better, Faster, Scalable, and Open Tabular Foundation ModelJingang QU, David Holzmüller, Gael Varoquaux, Marine Le MorvanICML 2026 · 被引用 85 次
- ALINE: Joint Amortization for Bayesian Inference and Active Data AcquisitionDaolang Huang, Xinyi Wen, Ayush Bharti, Samuel Kaski 等NeurIPS 2025 · 被引用 8 次
- In-Context Multi-Objective OptimizationXinyu Zhang, Conor Hassan, Julien Martinelli, Daolang Huang 等ICLR 2026 · 被引用 6 次
- Incremental Transformer Neural ProcessesPhilip Mortimer, Cristiana Diaconu, Tommy Rochussen, Bruno Mlodozeniec 等ICML 2026
- Revisiting Neural Processes via Fourier Transform and Volterra SeriesPeiman Mohseni, Nick Duffield, Raymond K WongICML 2026
它引用的顶会 Paper32
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
相关 Paper
- Autoregressive Conditional Neural ProcessesWessel P. Bruinsma, Stratis Markou, James Requeima, Andrew Y. K. Foong 等ICLR 2023 · 被引用 3 次
- Multi-Task Bayesian In-Context LearningQingyang Zhu, Eric Oermann, Kyunghyun ChoICML 2026
- Global Perception Based Autoregressive Neural ProcessesJinyang TaiICCV 2023 · 被引用 1 次
- Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence ModelingTung Nguyen, Aditya GroverICML 2022 · 被引用 148 次
- Distribution Transformers: Fast Approximate Bayesian Inference With On-The-Fly Prior AdaptationGeorge Whittle, Juliusz Ziomek, Jacob Rawling, Michael A OsborneICML 2026 · 被引用 13 次
