Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence Modeling
Tung Nguyen, Aditya Grover
摘要
Neural Processes (NPs) are a popular class of approaches for meta-learning. Similar to Gaussian Processes (GPs), NPs define distributions over functions and can estimate uncertainty in their predictions. However, unlike GPs, NPs and their variants suffer from underfitting and often have intractable likelihoods, which limit their applications in sequential decision making. We propose Transformer Neural Processes (TNPs), a new member of the NP family that casts uncertainty-aware meta learning as a sequence modeling problem. We learn TNPs via an autoregressive likelihood-based objective and instantiate it with a novel transformer-based architecture. The model architecture respects the inductive biases inherent to the problem structure, such as invariance to the observed data points and equivariance to the unobserved points. We further investigate knobs within the TNP framework that tradeoff expressivity of the decoding distribution with extra computation. Empirically, we show that TNPs achieve state-of-the-art performance on various benchmark problems, outperforming all previous NP variants on meta regression, image completion, contextual multi-armed bandits, and Bayesian optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper77
- What Can Transformers Learn In-Context? A Case Study of Simple Function ClassesShivam Garg, Dimitris Tsipras, Percy Liang, Gregory ValiantNeurIPS 2022 · 被引用 883 次
- How Transformers Learn Causal Structure with Gradient DescentEshaan Nichani, Alex Damian, Jason D. LeeICML 2024 · 被引用 117 次
- Diffusion Models for Black-Box OptimizationSiddarth Krishnamoorthy, Satvik Mehul Mashkaria, Aditya GroverICML 2023 · 被引用 94 次
- All-in-one simulation-based inferenceManuel Glöckler, Michael Deistler, Christian Dietrich Weilbach, Frank Wood 等ICML 2024 · 被引用 74 次
- LLM Processes: Numerical Predictive Distributions Conditioned on Natural LanguageJames Requeima, John Bronskill, Dami Choi, Richard E. Turner 等NeurIPS 2024 · 被引用 72 次
它引用的顶会 Paper14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee 等NeurIPS 2021 · 被引用 2,557 次
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
相关 Paper
- Latent Bottlenecked Attentive Neural ProcessesLeo Feng, Hossein Hajimirsadeghi, Yoshua Bengio, Mohamed Osama AhmedICLR 2023
- Practical Equivariances via Relational Conditional Neural ProcessesDaolang Huang, Manuel Haussmann, Ulpu Remes, S. T. John 等NeurIPS 2023 · 被引用 14 次
- Test Time Scaling for Neural ProcessesHyungi Lee, Moonseok Choi, Hyunsu Kim, Kyunghyun Cho 等NeurIPS 2025 · 被引用 1 次
- NPCL: Neural Processes for Uncertainty-Aware Continual LearningSaurav Jha, Dong Gong, He Zhao, Lina YaoNeurIPS 2023 · 被引用 27 次
- Learning to Generalize: An Information Perspective on Neural ProcessesHui Li, Huafeng Liu, Shuyang Lin, Jingyue Shi 等NeurIPS 2025
