Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence Modeling
Tung Nguyen, Aditya Grover
Abstract
Neural Processes (NPs) are a popular class of approaches for meta-learning. Similar to Gaussian Processes (GPs), NPs define distributions over functions and can estimate uncertainty in their predictions. However, unlike GPs, NPs and their variants suffer from underfitting and often have intractable likelihoods, which limit their applications in sequential decision making. We propose Transformer Neural Processes (TNPs), a new member of the NP family that casts uncertainty-aware meta learning as a sequence modeling problem. We learn TNPs via an autoregressive likelihood-based objective and instantiate it with a novel transformer-based architecture. The model architecture respects the inductive biases inherent to the problem structure, such as invariance to the observed data points and equivariance to the unobserved points. We further investigate knobs within the TNP framework that tradeoff expressivity of the decoding distribution with extra computation. Empirically, we show that TNPs achieve state-of-the-art performance on various benchmark problems, outperforming all previous NP variants on meta regression, image completion, contextual multi-armed bandits, and Bayesian optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e40a5ef1-66d4-4919-867b-e479144923bdCited by top-tier papers77
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 883 citations
- How Transformers Learn Causal Structure with Gradient DescentEshaan Nichani, Alex Damian, Jason D. LeeICML 2024 · 117 citations
- Diffusion Models for Black-Box OptimizationSiddarth Krishnamoorthy, Satvik Mehul Mashkaria, Aditya GroverICML 2023 · 94 citations
- All-in-one simulation-based inferenceManuel Glöckler, Michael Deistler, Christian Dietrich Weilbach, Frank Wood et al.ICML 2024 · 74 citations
- LLM Processes: Numerical Predictive Distributions Conditioned on Natural LanguageJames Requeima, John Bronskill, Dami Choi, Richard E. Turner et al.NeurIPS 2024 · 72 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
Related papers
- Latent Bottlenecked Attentive Neural ProcessesLeo Feng, Hossein Hajimirsadeghi, Yoshua Bengio, Mohamed Osama AhmedICLR 2023
- Practical Equivariances via Relational Conditional Neural ProcessesDaolang Huang, Manuel Haussmann, Ulpu Remes, S. T. John et al.NeurIPS 2023 · 14 citations
- Test Time Scaling for Neural ProcessesHyungi Lee, Moonseok Choi, Hyunsu Kim, Kyunghyun Cho et al.NeurIPS 2025 · 1 citation
- NPCL: Neural Processes for Uncertainty-Aware Continual LearningSaurav Jha, Dong Gong, He Zhao, Lina YaoNeurIPS 2023 · 27 citations
- Learning to Generalize: An Information Perspective on Neural ProcessesHui Li, Huafeng Liu, Shuyang Lin, Jingyue Shi et al.NeurIPS 2025
