Incremental Transformer Neural Processes
Philip Mortimer, Cristiana Diaconu, Tommy Rochussen, Bruno Mlodozeniec, Richard E Turner
摘要
Neural Processes (NPs), and specifically Transformer Neural Processes (TNPs), have demonstrated remarkable performance across tasks ranging from spatiotemporal forecasting to tabular data modelling. However, many of these applications are inherently sequential, involving continuous data streams such as real-time sensor readings or database updates. In such settings, models should support cheap, incremental updates rather than recomputing internal representations from scratch for every new observation—a capability existing TNP variants lack. Drawing inspiration from Large Language Models, we introduce the Incremental TNP (). By leveraging causal masking, Key-Value (KV) caching, and a data-efficient autoregressive training strategy, matches the predictive performance of standard TNPs while reducing the computational cost of updates from quadratic to linear time complexity. We empirically evaluate our model on a range of synthetic and real-world tasks, including tabular regression and temperature prediction. Our results show that, surprisingly, delivers performance comparable to—or better than—non-causal TNPs while unlocking orders-of-magnitude speedups for sequential inference. Finally, we assess the consistency of the model's updates---by adapting a metric of "implicit Bayesianness", we show that under a one-at-a-time streaming protocol, retains a prediction rule as implicitly Bayesian as standard non-causal TNPs, demonstrating that achieves the computational benefits of causal masking without sacrificing the consistency required for streaming inference.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka 等ICLR 2022 · 被引用 287 次
- Convolutional Conditional Neural ProcessesJonathan Gordon, Wessel P. Bruinsma, Andrew Y. K. Foong, James Requeima 等ICLR 2020 · 被引用 200 次
- Transformer Neural Processes: Uncertainty-Aware Meta Learning Via Sequence ModelingTung Nguyen, Aditya GroverICML 2022 · 被引用 148 次
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a SecondNoah Hollmann, Samuel Müller, Katharina Eggensperger, Frank HutterICLR 2023 · 被引用 96 次
相关 Paper
- Memory Efficient Neural Processes via Constant Memory Attention BlockLeo Feng, Frederick Tung, Hossein Hajimirsadeghi, Yoshua Bengio 等ICML 2024 · 被引用 8 次
- Efficient Autoregressive Inference for Transformer Probabilistic ModelsConor Hassan, Nasrulloh R. B. S. Loka, Cen-You Li, Daolang Huang 等ICLR 2026 · 被引用 5 次
- Variational Auto-Regressive Gaussian Processes for Continual LearningSanyam Kapoor, Theofanis Karaletsos, Thang D. BuiICML 2021 · 被引用 32 次
- Vector Quantized Bayesian Neural Network Inference for Data StreamsNamuk Park, Taekyu Lee, Songkuk KimAAAI 2021 · 被引用 11 次
- Streaming Visual Geometry TransformerDong Zhuo, Wenzhao Zheng, Jiahe Guo, Yuqi Wu 等ICLR 2026 · 被引用 109 次
