Semi-Parametric Inducing Point Networks and Neural Processes
Richa Rastogi, Yair Schiff, Alon Hacohen, Zhaozhi Li, Ian Lee, Yuntian Deng, Mert R. Sabuncu, Volodymyr Kuleshov
Abstract
We introduce semi-parametric inducing point networks (SPIN), a general-purpose architecture that can query the training set at inference time in a compute-efficient manner. Semi-parametric architectures are typically more compact than parametric models, but their computational complexity is often quadratic. In contrast, SPIN attains linear complexity via a cross-attention mechanism between datapoints inspired by inducing point methods. Querying large training sets can be particularly useful in meta-learning, as it unlocks additional training signal, but often exceeds the scaling limits of existing models. We use SPIN as the basis of the Inducing Point Neural Process, a probabilistic model which supports large contexts in meta-learning and achieves high accuracy where existing models fail. In our experiments, SPIN reduces memory requirements, improves accuracy across a range of meta-learning tasks, and improves state-of-the-art performance on an important practical problem, genotype imputation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7050c930-e513-422c-8001-ddd5c58d5e06Cited by top-tier papers6
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Diffusion Models With Learned Adaptive NoiseSubham S. Sahoo, Aaron Gokaslan, Christopher De Sa, Volodymyr KuleshovNeurIPS 2024 · 64 citations
- TuneTables: Context Optimization for Scalable Prior-Data Fitted NetworksBenjamin Feuer, Robin Schirrmeister, Valeriia Cherepanova, Chinmay Hegde et al.NeurIPS 2024 · 57 citations
- Theoretical Benefit and Limitation of Diffusion Language ModelGuhao Feng, Yihan Geng, Jian Guan, Wei Wu et al.NeurIPS 2025 · 52 citations
- Inducing Point Operator Transformer: A Flexible and Scalable Architecture for Solving PDEsSeungjun Lee, Taeil OhAAAI 2024 · 22 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
- Perceiver: General Perception with Iterative AttentionAndrew Jaegle, Felix Gimeno, Andy Brock, Oriol Vinyals et al.ICML 2021 · 1,399 citations
Related papers
- MARS: Meta-learning as Score Matching in the Function SpaceKrunoslav Lehman Pavasovic, Jonas Rothfuss, Andreas KrauseICLR 2023 · 1 citation
- Latent Bottlenecked Attentive Neural ProcessesLeo Feng, Hossein Hajimirsadeghi, Yoshua Bengio, Mohamed Osama AhmedICLR 2023
- Input Dependent Sparse Gaussian ProcessesBahram Jafrasteh, Carlos Villacampa-Calvo, Daniel Hernández-LobatoICML 2022 · 7 citations
- Accurate Bayesian Meta-Learning by Accurate Task Posterior InferenceMichael Volpp, Philipp Dahlinger, Philipp Becker, Christian Daniel et al.ICLR 2023
- Learning Large-scale Neural Fields via Context Pruned Meta-LearningJihoon Tack, Subin Kim, Sihyun Yu, Jaeho Lee et al.NeurIPS 2023 · 16 citations
