Learning under Model Misspecification: Applications to Variational and Ensemble methods
Andrés R. Masegosa
Abstract
Virtually any model we use in machine learning to make predictions does not perfectly represent reality. So, most of the learning happens under model misspecification. In this work, we present a novel analysis of the generalization performance of Bayesian model averaging under model misspecification and i.i.d. data using a new family of second-order PAC-Bayes bounds. This analysis shows, in simple and intuitive terms, that Bayesian model averaging provides suboptimal generalization performance when the model is misspecified. In consequence, we provide strong theoretical arguments showing that Bayesian methods are not optimal for learning predictive models, unless the model class is perfectly specified. Using novel second-order PAC-Bayes bounds, we derive a new family of Bayesian-like algorithms, which can be implemented as variational and ensemble methods. The output of these algorithms is a new posterior distribution, different from the Bayesian posterior, which induces a posterior predictive distribution with better generalization performance. Experiments with Bayesian neural networks illustrate these findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03b537f4-5f71-4bb3-965d-ae2637b2405eCited by top-tier papers19
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- How Good is the Bayes Posterior in Deep Neural Networks Really?Florian Wenzel, Kevin Roth, Bastiaan S. Veeling, Jakub Swiatkowski et al.ICML 2020 · 409 citations
- Deep Ensembles Work, But Are They Necessary?Taiga Abe, Estefany Kelly Buchanan, Geoff Pleiss, Richard S. Zemel et al.NeurIPS 2022 · 101 citations
- DICE: Diversity in Deep Ensembles via Conditional Redundancy Adversarial EstimationAlexandre Ramé, Matthieu CordICLR 2021 · 60 citations
- Dangers of Bayesian Model Averaging under Covariate ShiftPavel Izmailov, Patrick Nicholson, Sanae Lotfi, Andrew Gordon WilsonNeurIPS 2021 · 51 citations
Builds on4
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Parametric Gaussian Process RegressorsMartin Jankowiak, Geoff Pleiss, Jacob R. GardnerICML 2020 · 82 citations
- Second Order PAC-Bayesian Bounds for the Weighted Majority VoteAndrés R. Masegosa, Stephan Sloth Lorenzen, Christian Igel, Yevgeny SeldinNeurIPS 2020 · 48 citations
- Improved PAC-Bayesian Bounds for Linear RegressionVera Shalaeva, Alireza Fakhrizadeh Esfahani, Pascal Germain, Mihály PetreczkyAAAI 2020 · 20 citations
Related papers
- Loss function based second-order Jensen inequality and its application to particle variational inferenceFutoshi Futami, Tomoharu Iwata, Naonori Ueda, Issei Sato et al.NeurIPS 2021 · 5 citations
- PACOH: Bayes-Optimal Meta-Learning with PAC-GuaranteesJonas Rothfuss, Vincent Fortuin, Martin Josifoski, Andreas KrauseICML 2021 · 136 citations
- Flat Seeking Bayesian Neural NetworksVan-Anh Nguyen, Tung-Long Vuong, Hoang Phan, Thanh-Toan Do et al.NeurIPS 2023 · 14 citations
- More Flexible PAC-Bayesian Meta-Learning by Learning Learning AlgorithmsHossein Zakerinia, Amin Behjati, Christoph H. LampertICML 2024 · 11 citations
- Predictive variational inference: Learn the predictively optimal posterior distributionJinlin Lai, Antonio Linero, Yuling YaoICML 2026
