Repulsive Deep Ensembles are Bayesian
Francesco D'Angelo, Vincent Fortuin
摘要
Deep ensembles have recently gained popularity in the deep learning community for their conceptual simplicity and efficiency. However, maintaining functional diversity between ensemble members that are independently trained with gradient descent is challenging. This can lead to pathologies when adding more ensemble members, such as a saturation of the ensemble performance, which converges to the performance of a single model. Moreover, this does not only affect the quality of its predictions, but even more so the uncertainty estimates of the ensemble, and thus its performance on out-of-distribution data. We hypothesize that this limitation can be overcome by discouraging different ensemble members from collapsing to the same function. To this end, we introduce a kernelized repulsive term in the update rule of the deep ensembles. We show that this simple modification not only enforces and maintains diversity among the members but, even more importantly, transforms the maximum a posteriori inference into proper Bayesian inference. Namely, we show that the training dynamics of our proposed repulsive ensembles follow a Wasserstein gradient flow of the KL divergence with the true posterior. We study repulsive terms in weight and function space and empirically compare their performance to standard ensembles and Bayesian baselines on synthetic and real-world prediction tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- A Rigorous Link between Deep Ensembles and (Variational) Bayesian MethodsVeit David Wild, Sahra Ghalebikesabi, Dino Sejdinovic, Jeremias KnoblauchNeurIPS 2023 · 被引用 40 次
- Quantification of Uncertainty with Adversarial ModelsKajetan Schweighofer, Lukas Aichberger, Mykyta Ielanskyi, Günter Klambauer 等NeurIPS 2023 · 被引用 37 次
- Joint Training of Deep Ensembles Fails Due to Learner CollusionAlan Jeffares, Tennison Liu, Jonathan Crabbé, Mihaela van der SchaarNeurIPS 2023 · 被引用 34 次
- Generalized Variational Inference in Function Spaces: Gaussian Measures meet Bayesian Deep LearningVeit D. Wild, Robert Hu, Dino SejdinovicNeurIPS 2022 · 被引用 22 次
- Feature Kernel DistillationBobby He, Mete OzayICLR 2022 · 被引用 18 次
它引用的顶会 Paper17
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran 等NeurIPS 2020 · 被引用 604 次
- BatchEnsemble: an Alternative Approach to Efficient Ensemble and Lifelong LearningYeming Wen, Dustin Tran, Jimmy BaICLR 2020 · 被引用 569 次
- Continual learning with hypernetworksJohannes von Oswald, Christian Henning, João Sacramento, Benjamin F. GreweICLR 2020 · 被引用 412 次
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen 等ICLR 2020 · 被引用 292 次
相关 Paper
- Input-gradient space particle inference for neural network ensemblesTrung Q. Trinh, Markus Heinonen, Luigi Acerbi, Samuel KaskiICLR 2024 · 被引用 4 次
- Bayesian Deep Ensembles via the Neural Tangent KernelBobby He, Balaji Lakshminarayanan, Yee Whye TehNeurIPS 2020 · 被引用 136 次
- Diversity Matters When Learning From EnsemblesGiung Nam, Jongmin Yoon, Yoonho Lee, Juho LeeNeurIPS 2021 · 被引用 50 次
- Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKADavid Smerkous, Qinxun Bai, Fuxin LiNeurIPS 2024 · 被引用 3 次
- Bayesian Ensemble for Sequential Decision-MakingRui Liu, Enmin Zhao, Lu Wang, Yu Li 等ICLR 2026 · 被引用 2 次
