Adaptive kernel predictors from feature-learning infinite limits of neural networks
Clarissa Lauditi, Blake Bordelon, Cengiz Pehlevan
Abstract
Previous influential work showed that infinite width limits of neural networks in the lazy training regime are described by kernel machines. Here, we show that neural networks trained in the rich, feature learning infinite-width regime in two different settings are also described by kernel machines, but with data-dependent kernels. For both cases, we provide explicit expressions for the kernel predictors and prescriptions to numerically calculate them. To derive the first predictor, we study the large-width limit of feature-learning Bayesian networks, showing how feature learning leads to task-relevant adaptation of layer kernels and preactivation densities. The saddle point equations governing this limit result in a min-max optimization problem that defines the kernel predictor. To derive the second predictor, we study gradient flow training of randomly initialized networks trained with weight decay in the infinite-width limit using dynamical mean field theory (DMFT). The fixed point equations of the arising DMFT defines the task-adapted internal representations and the kernel predictor. We compare our kernel predictors to kernels derived from lazy regime and demonstrate that our adaptive kernels achieve lower test loss on benchmark datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38a7470c-c1a6-4dd0-af81-e9b9ff35daa6Cited by top-tier papers4
- A unified theory of feature learning in RNNs and DNNsJan Bauer, Kirsten Fischer, Moritz Helias, Agostina PalmigianoICML 2026 · 4 citations
- Dynamics of neural scaling laws in random feature regression with powerlaw-distributed kernel eigenvaluesJakob Kramp, Javed Lindner, Moritz HeliasICML 2026 · 2 citations
- Transfer Learning in Infinite Width Feature Learning NetworksClarissa Lauditi, Blake Bordelon, Cengiz PehlevanICLR 2026 · 2 citations
- To Grok Grokking: Provable Grokking in Ridge RegressionMingyue Xu, Gal Vardi, Itay SafranICML 2026
Builds on16
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 140 citations
- Neural Networks as Kernel Learners: The Silent Alignment EffectAlexander B. Atanasov, Blake Bordelon, Cengiz PehlevanICLR 2022 · 110 citations
- More Than a Toy: Random Matrix Models Predict How Real-World Neural Representations GeneralizeAlexander Wei, Wei Hu, Jacob SteinhardtICML 2022 · 90 citations
- Why bigger is not always better: on finite and infinite neural networksLaurence AitchisonICML 2020 · 59 citations
- Dynamics of Finite Width Kernel and Prediction Fluctuations in Mean Field Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2023 · 56 citations
Related papers
- The Influence of Learning Rule on Representation Dynamics in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanICLR 2023 · 7 citations
- Critical feature learning in deep neural networksKirsten Fischer, Javed Lindner, David Dahmen, Zohar Ringel et al.ICML 2024 · 15 citations
- A self consistent theory of Gaussian Processes captures feature learning effects in finite CNNsGadi Naveh, Zohar RingelNeurIPS 2021 · 38 citations
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Asymptotics of representation learning in finite Bayesian neural networksJacob A. Zavatone-Veth, Abdulkadir Canatar, Benjamin S. Ruben, Cengiz PehlevanNeurIPS 2021 · 45 citations
