Efficient and Scalable Bayesian Neural Nets with Rank-1 Factors
Michael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma, Jasper Snoek, Katherine A. Heller, Balaji Lakshminarayanan, Dustin Tran
Abstract
Bayesian neural networks (BNNs) demonstrate promising success in improving the robustness and uncertainty quantification of modern deep learning. However, they generally struggle with underfitting at scale and parameter efficiency. On the other hand, deep ensembles have emerged as alternatives for uncertainty quantification that, while outperforming BNNs on certain problems, also suffer from efficiency issues. It remains unclear how to combine the strengths of these two approaches and remediate their common issues. To tackle this challenge, we propose a rank-1 parameterization of BNNs, where each weight matrix involves only a distribution on a rank-1 subspace. We also revisit the use of mixture approximate posteriors to capture multiple modes, where unlike typical mixtures, this approach admits a significantly smaller memory increase (e.g., only a 0.4% increase for a ResNet-50 mixture of size 10). We perform a systematic empirical study on the choices of prior, variational posterior, and methods to improve training. For ResNet-50 on ImageNet, Wide ResNet 28-10 on CIFAR-10/100, and an RNN on MIMIC-III, rank-1 BNNs achieve state-of-the-art performance across log-likelihood, accuracy, and calibration on the test sets and outof-distribution variants. 1 * Equal contribution † Work completed as a Google AI Resident.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 354ab3f4-d45b-41ea-86c5-cafe40d6c7ceCited by top-tier papers83
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran et al.NeurIPS 2020 · 604 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 263 citations
- Training independent subnetworks for robust predictionMarton Havasi, Rodolphe Jenatton, Stanislav Fort, Jeremiah Zhe Liu et al.ICLR 2021 · 235 citations
- Bayesian Neural Network Priors RevisitedVincent Fortuin, Adrià Garriga-Alonso, Sebastian W. Ober, Florian Wenzel et al.ICLR 2022 · 162 citations
Builds on2
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen et al.ICLR 2020 · 292 citations
- The k-tied Normal Distribution: A Compact Parameterization of Gaussian Mean Field Posteriors in Bayesian Neural NetworksJakub Swiatkowski, Kevin Roth, Bastiaan S. Veeling, Linh Tran et al.ICML 2020 · 52 citations
Related papers
- Sparse Uncertainty Representation in Deep Learning with Inducing WeightsHippolyt Ritter, Martin Kukla, Cheng Zhang, Yingzhen LiNeurIPS 2021 · 23 citations
- Make Me a BNN: A Simple Strategy for Estimating Bayesian Uncertainty from Pre-trained ModelsGianni Franchi, Olivier Laurent, Maxence Leguéry, Andrei Bursuc et al.CVPR 2024
- Bayesian Low-Rank Learning (Bella): A Practical Approach to Bayesian Neural NetworksBao Gia Doan, Afshar Shamsi, Xiao-Yu Guo, Arash Mohammadi et al.AAAI 2025 · 1 citation
- Flat Seeking Bayesian Neural NetworksVan-Anh Nguyen, Tung-Long Vuong, Hoang Phan, Thanh-Toan Do et al.NeurIPS 2023 · 14 citations
- On the Expressiveness of Approximate Inference in Bayesian Neural NetworksAndrew Y. K. Foong, David R. Burt, Yingzhen Li, Richard E. TurnerNeurIPS 2020 · 142 citations
