Squared families are useful conjugate priors
Russell Tsuchida, Jiawei Liu, Cheng Soon Ong, Dino Sejdinovic
Abstract
Squared families of probability distributions have been studied and applied in numerous machine learning contexts. Typically, they appear as likelihoods, where their advantageous computational, geometric and statistical properties are exploited for fast estimation algorithms, representational properties and statistical guarantees. Here, we investigate the use of squared families as prior beliefs in Bayesian inference. We find that they can form helpful conjugate families, often allowing for closed-form and tractable Bayesian inference and marginal likelihoods. We apply such conjugate families to Bayesian regression in feature space using end-to-end learnable neural network features. Such a setting allows for a rich multi-modal alternative to Gaussian processes with neural network features, often called deep kernel learning. We demonstrate our method on few shot learning, outperforming existing neural methods based on Gaussian processes and normalising flows. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Bayesian Meta-Learning for the Few-Shot Setting via Deep KernelsMassimiliano Patacchiola, Jack Turner, Elliot J. Crowley, Michael F. P. O'Boyle et al.NeurIPS 2020 · 167 citations
- Non-parametric Models for Non-negative FunctionsUlysse Marteau-Ferey, Francis R. Bach, Alessandro RudiNeurIPS 2020 · 65 citations
- How to Turn Your Knowledge Graph Embeddings into Generative ModelsLorenzo Loconte, Nicola Di Mauro, Robert Peharz, Antonio VergariNeurIPS 2023 · 30 citations
- PSD Representations for Effective Probability ModelsAlessandro Rudi, Carlo CilibertoNeurIPS 2021 · 28 citations
- Fast Neural Kernel Embeddings for General ActivationsInsu Han, Amir Zandieh, Jaehoon Lee, Roman Novak et al.NeurIPS 2022 · 26 citations
Related papers
- Squared Neural Families: A New Class of Tractable Density ModelsRussell Tsuchida, Cheng Soon Ong, Dino SejdinovicNeurIPS 2023 · 15 citations
- Revisiting Logistic-softmax Likelihood in Bayesian Meta-Learning for Few-Shot ClassificationTianjun Ke, Haoqun Cao, Zenan Ling, Feng ZhouNeurIPS 2023 · 17 citations
- Transformers Can Do Bayesian InferenceSamuel Müller, Noah Hollmann, Sebastian Pineda-Arango, Josif Grabocka et al.ICLR 2022 · 287 citations
- Non-Gaussian Gaussian Processes for Few-Shot RegressionMarcin Sendera, Jacek Tabor, Aleksandra Nowak, Andrzej Bedychaj et al.NeurIPS 2021 · 23 citations
- Learning to Learn Dense Gaussian Processes for Few-Shot LearningZe Wang, Zichen Miao, Xiantong Zhen, Qiang QiuNeurIPS 2021 · 32 citations
