Bayesian Model Selection, the Marginal Likelihood, and Generalization
Sanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum, Andrew Gordon Wilson
Abstract
How do we compare between hypotheses that are entirely consistent with observations? The marginal likelihood (aka Bayesian evidence), which represents the probability of generating our observations from a prior, provides a distinctive approach to this foundational question, automatically encoding Occam's razor. Although it has been observed that the marginal likelihood can overfit and is sensitive to prior assumptions, its limitations for hyperparameter learning and discrete model comparison have not been thoroughly investigated. We first revisit the appealing properties of the marginal likelihood for learning constraints and hypothesis testing. We then highlight the conceptual and practical issues in using the marginal likelihood as a proxy for generalization. Namely, we show how marginal likelihood can be negatively correlated with generalization, with implications for neural architecture search, and can lead to both underfitting and overfitting in hyperparameter learning. We also re-examine the connection between the marginal likelihood and PAC-Bayes bounds and use this connection to further elucidate the shortcomings of the marginal likelihood for model selection. We provide a partial remedy through a conditional marginal likelihood, which we show is more aligned with generalization, and practically valuable for large-scale hyperparameter learning, such as in deep kernel learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb5a5a9d-9411-4dd0-8a78-1a5394ea2056Cited by top-tier papers20
- A Study of Bayesian Neural Network Surrogates for Bayesian OptimizationYucen Lily Li, Tim G. J. Rudner, Andrew Gordon WilsonICLR 2024 · 59 citations
- GFlowOut: Dropout with Generative Flow NetworksDianbo Liu, Moksh Jain, Bonaventure F. P. Dossou, Qianli Shen et al.ICML 2023 · 27 citations
- Posterior Refinement Improves Sample Efficiency in Bayesian Neural NetworksAgustinus Kristiadi, Runa Eschenhagen, Philipp HennigNeurIPS 2022 · 17 citations
- PAC-Bayes-Chernoff bounds for unbounded lossesIoar Casado, Luis A. Ortega Andrés, Aritz Pérez, Andrés R. MasegosaNeurIPS 2024 · 14 citations
- Robust Gaussian Processes via Relevance PursuitSebastian Ament, Elizabeth Santorella, David Eriksson, Ben Letham et al.NeurIPS 2024 · 12 citations
Builds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen et al.NeurIPS 2021 · 508 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
Related papers
- Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks using the Marginal LikelihoodRayen Dhahri, Alexander Immer, Bertrand Charpentier, Stephan Günnemann et al.NeurIPS 2024 · 10 citations
- A Bayesian Perspective on Training Speed and Model SelectionClare Lyle, Lisa Schut, Binxin Ru, Yarin Gal et al.NeurIPS 2020 · 26 citations
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch et al.ICML 2021 · 130 citations
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski et al.NeurIPS 2022 · 98 citations
- Generalization Guarantees for Neural Architecture Search with Train-Validation SplitSamet Oymak, Mingchen Li, Mahdi SoltanolkotabiICML 2021 · 20 citations
