A theory of representation learning gives a deep generalisation of kernel methods
Adam X. Yang, Maxime Robeyns, Edward Milsom, Ben Anson, Nandi Schoots, Laurence Aitchison
Abstract
The successes of modern deep machine learning methods are founded on their ability to transform inputs across multiple layers to build good high-level representations. It is therefore critical to understand this process of representation learning. However, standard theoretical approaches (formally NNGPs) involving infinite width limits eliminate representation learning. We therefore develop a new infinite width limit, the Bayesian representation learning limit, that exhibits representation learning mirroring that in finite-width models, yet at the same time, retains some of the simplicity of standard infinite-width limits. In particular, we show that Deep Gaussian processes (DGPs) in the Bayesian representation learning limit have exactly multivariate Gaussian posteriors, and the posterior covariances can be obtained by optimizing an interpretable objective combining a log-likelihood to improve performance with a series of KL-divergences which keep the posteriors close to the prior. We confirm these results experimentally in wide but finite DGPs. Next, we introduce the possibility of using this limit and objective as a flexible, deep generalisation of kernel methods, that we call deep kernel machines (DKMs). Like most naive kernel methods, DKMs scale cubically in the number of datapoints. We therefore use methods from the Gaussian process inducing point literature to develop a sparse DKM that scales linearly in the number of datapoints. Finally, we extend these approaches to NNs (which have non-Gaussian posteriors) in the Appendices.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30ba8aca-a40a-48be-a3c2-375ccfcc3d9bCited by top-tier papers10
- Dynamics of Finite Width Kernel and Prediction Fluctuations in Mean Field Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2023 · 56 citations
- DePT: Decomposed Prompt Tuning for Parameter-Efficient Fine-tuningZhengxiang Shi, Aldo LipaniICLR 2024 · 46 citations
- The Empirical Impact of Neural Parameter Symmetries, or Lack ThereofDerek Lim, Theo (Moe) Putterman, Robin Walters, Haggai Maron et al.NeurIPS 2024 · 25 citations
- Critical feature learning in deep neural networksKirsten Fischer, Javed Lindner, David Dahmen, Zohar Ringel et al.ICML 2024 · 15 citations
- Convolutional Deep Kernel MachinesEdward Milsom, Ben Anson, Laurence AitchisonICLR 2024 · 6 citations
Builds on10
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Asymptotics of Wide Networks from Feynman DiagramsEthan Dyer, Guy Gur-AriICLR 2020 · 127 citations
- Why bigger is not always better: on finite and infinite neural networksLaurence AitchisonICML 2020 · 59 citations
- Asymptotics of representation learning in finite Bayesian neural networksJacob A. Zavatone-Veth, Abdulkadir Canatar, Benjamin S. Ruben, Cengiz PehlevanNeurIPS 2021 · 45 citations
- Deep Kernel ProcessesLaurence Aitchison, Adam X. Yang, Sebastian W. OberICML 2021 · 44 citations
Related papers
- Bayesian Deep Ensembles via the Neural Tangent KernelBobby He, Balaji Lakshminarayanan, Yee Whye TehNeurIPS 2020 · 136 citations
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process PerspectiveGeoff Pleiss, John P. CunninghamNeurIPS 2021 · 35 citations
- A variational approximate posterior for the deep Wishart processSebastian W. Ober, Laurence AitchisonNeurIPS 2021 · 11 citations
- Deep Kernel Posterior Learning under Infinite Variance Prior WeightsJorge Loría, Anindya BhadraICLR 2025
- On Neural Networks as Infinite Tree-Structured Probabilistic Graphical ModelsBoyao Li, Alexander Thomson, Houssam Nassif, Matthew Engelhard et al.NeurIPS 2024 · 2 citations
