Deep Kernel Processes
Laurence Aitchison, Adam X. Yang, Sebastian W. Ober
Abstract
We define deep kernel processes in which positive definite Gram matrices are progressively transformed by nonlinear kernel functions and by sampling from (inverse) Wishart distributions. Remarkably, we find that deep Gaussian processes (DGPs), Bayesian neural networks (BNNs), infinite BNNs, and infinite BNNs with bottlenecks can all be written as deep kernel processes. For DGPs the equivalence arises because the Gram matrix formed by the inner product of features is Wishart distributed, and as we show, standard isotropic kernels can be written entirely in terms of this Gram matrix -- we do not need knowledge of the underlying features. We define a tractable deep kernel process, the deep inverse Wishart process, and give a doubly-stochastic inducing-point variational inference scheme that operates on the Gram matrices, not on the features, as in DGPs. We show that the deep inverse Wishart process gives superior performance to DGPs and infinite BNNs on standard fully-connected baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Bayesian Neural Network Priors RevisitedVincent Fortuin, Adrià Garriga-Alonso, Sebastian W. Ober, Florian Wenzel et al.ICLR 2022 · 162 citations
- Bayesian Low-rank Adaptation for Large Language ModelsAdam X. Yang, Maxime Robeyns, Xi Wang, Laurence AitchisonICLR 2024 · 111 citations
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processesSebastian W. Ober, Laurence AitchisonICML 2021 · 65 citations
- Precise characterization of the prior predictive distribution of deep ReLU networksLorenzo Noci, Gregor Bachmann, Kevin Roth, Sebastian Nowozin et al.NeurIPS 2021 · 36 citations
Builds on3
- Global inducing point variational posteriors for Bayesian neural networks and deep Gaussian processesSebastian W. Ober, Laurence AitchisonICML 2021 · 65 citations
- Why bigger is not always better: on finite and infinite neural networksLaurence AitchisonICML 2020 · 59 citations
- Stochastic Differential Equations with Variational Wishart DiffusionsMartin Jørgensen, Marc Peter Deisenroth, Hugh SalimbeniICML 2020 · 8 citations
Related papers
- A variational approximate posterior for the deep Wishart processSebastian W. Ober, Laurence AitchisonNeurIPS 2021 · 11 citations
- Deep Kernel Posterior Learning under Infinite Variance Prior WeightsJorge Loría, Anindya BhadraICLR 2025
- A theory of representation learning gives a deep generalisation of kernel methodsAdam X. Yang, Maxime Robeyns, Edward Milsom, Ben Anson et al.ICML 2023 · 15 citations
- On Neural Networks as Infinite Tree-Structured Probabilistic Graphical ModelsBoyao Li, Alexander Thomson, Houssam Nassif, Matthew Engelhard et al.NeurIPS 2024 · 2 citations
- Bayesian Deep Ensembles via the Neural Tangent KernelBobby He, Balaji Lakshminarayanan, Yee Whye TehNeurIPS 2020 · 136 citations
