Scalable Marginal Likelihood Estimation for Model Selection in Deep Learning
Alexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch, Mohammad Emtiyaz Khan
Abstract
Marginal-likelihood based model-selection, even though promising, is rarely used in deep learning due to estimation difficulties. Instead, most approaches rely on validation data, which may not be readily available. In this work, we present a scalable marginal-likelihood estimation method to select both hyperparameters and network architectures, based on the training data alone. Some hyperparameters can be estimated online during training, simplifying the procedure. Our marginal-likelihood estimate is based on Laplace's method and Gauss-Newton approximations to the Hessian, and it outperforms cross-validation and manual-tuning on standard regression and image classification datasets, especially in terms of calibration and out-of-distribution detection. Our work shows that marginal likelihoods can improve generalization and be useful when validation data is unavailable (e.g., in nonstationary settings).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5887c253-ceea-4239-b02d-df3d5e1d7296Cited by top-tier papers44
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
- Bayesian Neural Network Priors RevisitedVincent Fortuin, Adrià Garriga-Alonso, Sebastian W. Ober, Florian Wenzel et al.ICLR 2022 · 162 citations
- Repulsive Deep Ensembles are BayesianFrancesco D'Angelo, Vincent FortuinNeurIPS 2021 · 141 citations
- Bayesian Deep Learning via Subnetwork InferenceErik A. Daxberger, Eric T. Nalisnick, James Urquhart Allingham, Javier Antorán et al.ICML 2021 · 108 citations
- On Uncertainty, Tempering, and Data Augmentation in Bayesian ClassificationSanyam Kapoor, Wesley J. Maddox, Pavel Izmailov, Andrew Gordon WilsonNeurIPS 2022 · 64 citations
Builds on4
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 344 citations
- A Bayesian Perspective on Training Speed and Model SelectionClare Lyle, Lisa Schut, Binxin Ru, Yarin Gal et al.NeurIPS 2020 · 26 citations
Related papers
- Stochastic Marginal Likelihood Gradients using Neural Tangent KernelsAlexander Immer, Tycho F. A. van der Ouderaa, Mark van der Wilk, Gunnar Rätsch et al.ICML 2023 · 17 citations
- Invariance Learning in Deep Neural Networks with Differentiable Laplace ApproximationsAlexander Immer, Tycho F. A. van der Ouderaa, Gunnar Rätsch, Vincent Fortuin et al.NeurIPS 2022 · 56 citations
- Hyperparameter Optimization through Neural Network PartitioningBruno Mlodozeniec, Matthias Reisser, Christos LouizosICLR 2023
- Adapting the Linearised Laplace Model Evidence for Modern Deep LearningJavier Antorán, David Janz, James Urquhart Allingham, Erik A. Daxberger et al.ICML 2022 · 36 citations
- Effective Bayesian Heteroscedastic Regression with Deep Neural NetworksAlexander Immer, Emanuele Palumbo, Alexander Marx, Julia E. VogtNeurIPS 2023 · 34 citations
