Repulsive Deep Ensembles are Bayesian
Francesco D'Angelo, Vincent Fortuin
Abstract
Deep ensembles have recently gained popularity in the deep learning community for their conceptual simplicity and efficiency. However, maintaining functional diversity between ensemble members that are independently trained with gradient descent is challenging. This can lead to pathologies when adding more ensemble members, such as a saturation of the ensemble performance, which converges to the performance of a single model. Moreover, this does not only affect the quality of its predictions, but even more so the uncertainty estimates of the ensemble, and thus its performance on out-of-distribution data. We hypothesize that this limitation can be overcome by discouraging different ensemble members from collapsing to the same function. To this end, we introduce a kernelized repulsive term in the update rule of the deep ensembles. We show that this simple modification not only enforces and maintains diversity among the members but, even more importantly, transforms the maximum a posteriori inference into proper Bayesian inference. Namely, we show that the training dynamics of our proposed repulsive ensembles follow a Wasserstein gradient flow of the KL divergence with the true posterior. We study repulsive terms in weight and function space and empirically compare their performance to standard ensembles and Bayesian baselines on synthetic and real-world prediction tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2436bf36-38dc-45df-996b-fe5e1ae388feCited by top-tier papers31
- A Rigorous Link between Deep Ensembles and (Variational) Bayesian MethodsVeit David Wild, Sahra Ghalebikesabi, Dino Sejdinovic, Jeremias KnoblauchNeurIPS 2023 · 40 citations
- Quantification of Uncertainty with Adversarial ModelsKajetan Schweighofer, Lukas Aichberger, Mykyta Ielanskyi, Günter Klambauer et al.NeurIPS 2023 · 37 citations
- Joint Training of Deep Ensembles Fails Due to Learner CollusionAlan Jeffares, Tennison Liu, Jonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 34 citations
- Generalized Variational Inference in Function Spaces: Gaussian Measures meet Bayesian Deep LearningVeit D. Wild, Robert Hu, Dino SejdinovicNeurIPS 2022 · 22 citations
- Feature Kernel DistillationBobby He, Mete OzayICLR 2022 · 18 citations
Builds on17
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran et al.NeurIPS 2020 · 604 citations
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 569 citations
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 412 citations
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen et al.ICLR 2020 · 292 citations
Related papers
- Input-gradient space particle inference for neural network ensemblesTrung Q. Trinh, Markus Heinonen, Luigi Acerbi, Samuel KaskiICLR 2024 · 4 citations
- Bayesian Deep Ensembles via the Neural Tangent KernelBobby He, Balaji Lakshminarayanan, Yee Whye TehNeurIPS 2020 · 136 citations
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 50 citations
- Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKADavid Smerkous, Qinxun Bai, Fuxin LiNeurIPS 2024 · 3 citations
- Bayesian Ensemble for Sequential Decision-MakingRui Liu, Enmin Zhao, Lu Wang, Yu Li et al.ICLR 2026 · 2 citations
