Stochastic Marginal Likelihood Gradients using Neural Tangent Kernels
Alexander Immer, Tycho F. A. van der Ouderaa, Mark van der Wilk, Gunnar Rätsch, Bernhard Schölkopf
摘要
Selecting hyperparameters in deep learning greatly impacts its effectiveness but requires manual effort and expertise. Recent works show that Bayesian model selection with Laplace approximations can allow to optimize such hyperparameters just like standard neural network parameters using gradients and on the training data. However, estimating a single hyperparameter gradient requires a pass through the entire dataset, limiting the scalability of such algorithms. In this work, we overcome this issue by introducing lower bounds to the linearized Laplace approximation of the marginal likelihood. In contrast to previous estimators, these bounds are amenable to stochastic-gradient-based optimization and allow to trade off estimation accuracy against computational complexity. We derive them using the function-space form of the linearized Laplace, which can be estimated using the neural tangent kernel. Experimentally, we show that the estimators can significantly accelerate gradient-based hyperparameter optimization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Effective Bayesian Heteroscedastic Regression with Deep Neural NetworksAlexander Immer, Emanuele Palumbo, Alexander Marx, Julia E. VogtNeurIPS 2023 · 被引用 34 次
- Learning Layer-wise Equivariances Automatically using GradientsTycho F. A. van der Ouderaa, Alexander Immer, Mark van der WilkNeurIPS 2023 · 被引用 28 次
- Improving Neural Additive Models with Bayesian PrinciplesKouroche Bouchiat, Alexander Immer, Hugo Yèche, Gunnar Rätsch 等ICML 2024 · 被引用 17 次
- A Generative Model of Symmetry TransformationsJames Urquhart Allingham, Bruno Mlodozeniec, Shreyas Padhy, Javier Antorán 等NeurIPS 2024 · 被引用 16 次
- Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks using the Marginal LikelihoodRayen Dhahri, Alexander Immer, Bertrand Charpentier, Stephan Günnemann 等NeurIPS 2024 · 被引用 10 次
它引用的顶会 Paper14
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 等ICLR 2020 · 被引用 254 次
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch 等ICML 2021 · 被引用 130 次
- Approximately Equivariant Networks for Imperfectly Symmetric DynamicsRui Wang, Robin Walters, Rose YuICML 2022 · 被引用 111 次
- Fast Finite Width Neural Tangent KernelRoman Novak, Jascha Sohl-Dickstein, Samuel S. SchoenholzICML 2022 · 被引用 72 次
相关 Paper
- Adapting the Linearised Laplace Model Evidence for Modern Deep LearningJavier Antorán, David Janz, James Urquhart Allingham, Erik A. Daxberger 等ICML 2022 · 被引用 36 次
- Accelerated Linearized Laplace Approximation for Bayesian Deep LearningZhijie Deng, Feng Zhou, Jun ZhuNeurIPS 2022 · 被引用 36 次
- Sample complexity of data-driven tuning of model hyperparameters in neural networks with structured parameter-dependent dual functionMaria-Florina Balcan, Anh Nguyen, Dravyansh SharmaNeurIPS 2025 · 被引用 14 次
- Hamiltonian Monte Carlo using an adjoint-differentiated Laplace approximation: Bayesian inference for latent Gaussian models and beyondCharles C. Margossian, Aki Vehtari, Daniel Simpson, Raj AgrawalNeurIPS 2020 · 被引用 30 次
- Invariance Learning in Deep Neural Networks with Differentiable Laplace ApproximationsAlexander Immer, Tycho F. A. van der Ouderaa, Gunnar Rätsch, Vincent Fortuin 等NeurIPS 2022 · 被引用 56 次
