Deep Kernel Posterior Learning under Infinite Variance Prior Weights
Jorge Loría, Anindya Bhadra
Abstract
Neal (1996) proved that infinitely wide shallow Bayesian neural networks (BNN) converge to Gaussian processes (GP), when the network weights have bounded prior variance. Cho & Saul (2009) provided a useful recursive formula for deep kernel processes for relating the covariance kernel of each layer to the layer immediately below. Moreover, they worked out the form of the layer-wise covariance kernel in an explicit manner for several common activation functions, including the ReLU. Subsequent works have made the connection between these two works, and provided useful results on the covariance kernel of a deep GP arising as wide limits of various deep Bayesian network architectures. However, recent works, including Aitchison et al. (2021), have highlighted that the covariance kernels obtained in this manner are deterministic and hence, precludes any possibility of representation learning, which amounts to learning a non-degenerate posterior of a random kernel given the data. To address this, they propose adding artificial noise to the kernel to retain stochasticity, and develop deep kernel Wishart and inverse Wishart processes. Nonetheless, this artificial noise injection could be critiqued in that it would not naturally emerge in a classic BNN architecture under an infinite-width limit. To address this, we show that a Bayesian deep neural network, where each layer width approaches infinity, and all network weights are elliptically distributed with infinite variance, converges to a process with -stable marginals in each layer that has a conditionally Gaussian representation. These conditional random covariance kernels could be recursively linked in the manner of Cho & Saul (2009), even though marginally the process exhibits stable behavior, and hence covariances are not even necessarily defined. We also provide useful generalizations of the recent results of Loría & Bhadra (2024) on shallow networks to multi-layer networks, and remedy the prohibitive computational burden of their approach. The computational and statistical benefits over competing approaches stand out in simulations and in demonstrations on benchmark data sets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 64f85a0c-dda9-42f7-82c2-16cbcd3aa2d1Builds on4
- Why bigger is not always better: on finite and infinite neural networksLaurence AitchisonICML 2020 · 59 citations
- Deep Kernel ProcessesLaurence Aitchison, Adam X. Yang, Sebastian W. OberICML 2021 · 44 citations
- A theory of representation learning gives a deep generalisation of kernel methodsAdam X. Yang, Maxime Robeyns, Edward Milsom, Ben Anson et al.ICML 2023 · 15 citations
- Avoiding Kernel Fixed Points: Computing with ELU and GELU Infinite NetworksRussell Tsuchida, Tim Pearce, Christopher van der Heide, Fred Roosta et al.AAAI 2021 · 10 citations
Related papers
- Critical feature learning in deep neural networksKirsten Fischer, Javed Lindner, David Dahmen, Zohar Ringel et al.ICML 2024 · 15 citations
- Beyond IID weights: sparse and low-rank deep Neural Networks are also Gaussian ProcessesThiziri Nait Saada, Alireza Naderi, Jared TannerICLR 2024 · 2 citations
- A variational approximate posterior for the deep Wishart processSebastian W. Ober, Laurence AitchisonNeurIPS 2021 · 11 citations
- Large-width functional asymptotics for deep Gaussian neural networksDaniele Bracale, Stefano Favaro, Sandra Fortini, Stefano PeluchettiICLR 2021 · 2 citations
- An Infinite-Feature Extension for Bayesian ReLU Nets That Fixes Their Asymptotic OverconfidenceAgustinus Kristiadi, Matthias Hein, Philipp HennigNeurIPS 2021 · 10 citations
