Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels
Alexander Immer, Tycho F. A. van der Ouderaa, Mark van der Wilk, Gunnar Rätsch, Bernhard Schölkopf
Abstract
Selecting hyperparameters in deep learning greatly impacts its effectiveness but requires manual effort and expertise. Recent works show that Bayesian model selection with Laplace approximations can allow to optimize such hyperparameters just like standard neural network parameters using gradients and on the training data. However, estimating a single hyperparameter gradient requires a pass through the entire dataset, limiting the scalability of such algorithms. In this work, we overcome this issue by introducing lower bounds to the linearized Laplace approximation of the marginal likelihood. In contrast to previous estimators, these bounds are amenable to stochastic-gradient-based optimization and allow to trade off estimation accuracy against computational complexity. We derive them using the function-space form of the linearized Laplace, which can be estimated using the neural tangent kernel. Experimentally, we show that the estimators can significantly accelerate gradient-based hyperparameter optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Effective Bayesian Heteroscedastic Regression with Deep Neural NetworksAlexander Immer, Emanuele Palumbo, Alexander Marx, Julia E. VogtNeurIPS 2023 · 34 citations
- Learning Layer-wise Equivariances Automatically using GradientsTycho F. A. van der Ouderaa, Alexander Immer, Mark van der WilkNeurIPS 2023 · 28 citations
- Improving Neural Additive Models with Bayesian PrinciplesKouroche Bouchiat, Alexander Immer, Hugo Yèche, Gunnar Rätsch et al.ICML 2024 · 17 citations
- A Generative Model of Symmetry TransformationsJames Urquhart Allingham, Bruno Mlodozeniec, Shreyas Padhy, Javier Antorán et al.NeurIPS 2024 · 16 citations
- Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks using the Marginal LikelihoodRayen Dhahri, Alexander Immer, Bertrand Charpentier, Stephan Günnemann et al.NeurIPS 2024 · 10 citations
Builds on14
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee et al.ICLR 2020 · 254 citations
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch et al.ICML 2021 · 130 citations
- Approximately Equivariant Networks for Imperfectly Symmetric DynamicsRui Wang, Robin Walters, Rose YuICML 2022 · 111 citations
- Fast Finite Width Neural Tangent KernelRoman Novak, Jascha Sohl-Dickstein, Samuel S. SchoenholzICML 2022 · 72 citations
Related papers
- Adapting the Linearised Laplace Model Evidence for Modern Deep LearningJavier Antorán, David Janz, James Urquhart Allingham, Erik A. Daxberger et al.ICML 2022 · 36 citations
- Accelerated Linearized Laplace Approximation for Bayesian Deep LearningZhijie Deng, Feng Zhou, Jun ZhuNeurIPS 2022 · 36 citations
- Sample complexity of data-driven tuning of model hyperparameters in neural networks with structured parameter-dependent dual functionMaria-Florina Balcan, Anh Nguyen, Dravyansh SharmaNeurIPS 2025 · 14 citations
- Hamiltonian Monte Carlo using an adjoint-differentiated Laplace approximation: Bayesian inference for latent Gaussian models and beyondCharles C. Margossian, Aki Vehtari, Daniel Simpson, Raj AgrawalNeurIPS 2020 · 30 citations
- Invariance Learning in Deep Neural Networks with Differentiable Laplace ApproximationsAlexander Immer, Tycho F. A. van der Ouderaa, Gunnar Rätsch, Vincent Fortuin et al.NeurIPS 2022 · 56 citations
