A variational approximate posterior for the deep Wishart process
Sebastian W. Ober, Laurence Aitchison
Abstract
Recent work introduced deep kernel processes as an entirely kernel-based alternative to NNs (Aitchison et al. 2020) . Deep kernel processes flexibly learn good top-layer representations by alternately sampling the kernel from a distribution over positive semi-definite matrices and performing nonlinear transformations. A particular deep kernel process, the deep Wishart process (DWP), is of particular interest because its prior can be made equivalent to deep Gaussian process (DGP) priors for kernels that can be expressed entirely in terms of Gram matrices. However, inference in DWPs has not yet been possible due to the lack of sufficiently flexible distributions over positive semi-definite matrices. Here, we give a novel approach to obtaining flexible distributions over positive semi-definite matrices by generalising the Bartlett decomposition of the Wishart probability density. We use this new distribution to develop an approximate posterior for the DWP that includes dependency across layers. We develop a doubly-stochastic inducing-point inference scheme for the DWP and show experimentally that inference in the DWP can improve performance over doing inference in a DGP with the equivalent prior.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02b450f4-ea8a-4504-9934-b9c4e051b0aeCited by top-tier papers2
- A theory of representation learning gives a deep generalisation of kernel methodsAdam X. Yang, Maxime Robeyns, Edward Milsom, Ben Anson et al.ICML 2023 · 15 citations
- Stochastic Kernel Regularisation Improves Generalisation in Deep Kernel MachinesEdward Milsom, Ben Anson, Laurence AitchisonNeurIPS 2024 · 1 citation
Builds on3
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processesSebastian W. Ober, Laurence AitchisonICML 2021 · 65 citations
- Why bigger is not always better: on finite and infinite neural networksLaurence AitchisonICML 2020 · 59 citations
- Stochastic Differential Equations with Variational Wishart DiffusionsMartin Jørgensen, Marc Peter Deisenroth, Hugh SalimbeniICML 2020 · 8 citations
Related papers
- Deep Kernel ProcessesLaurence Aitchison, Adam X. Yang, Sebastian W. OberICML 2021 · 44 citations
- Deep Kernel Posterior Learning under Infinite Variance Prior WeightsJorge Loría, Anindya BhadraICLR 2025
- Bayesian Deep Ensembles via the Neural Tangent KernelBobby He, Balaji Lakshminarayanan, Yee Whye TehNeurIPS 2020 · 136 citations
- Convolutional Deep Kernel MachinesEdward Milsom, Ben Anson, Laurence AitchisonICLR 2024 · 6 citations
- On Neural Networks as Infinite Tree-Structured Probabilistic Graphical ModelsBoyao Li, Alexander Thomson, Houssam Nassif, Matthew Engelhard et al.NeurIPS 2024 · 2 citations
