Training Bayesian Neural Networks with Sparse Subspace Variational Inference
Junbo Li, Zichen Miao, Qiang Qiu, Ruqi Zhang
Abstract
Bayesian neural networks (BNNs) offer uncertainty quantification but come with the downside of substantially increased training and inference costs. Sparse BNNs have been investigated for efficient inference, typically by either slowly introducing sparsity throughout the training or by post-training compression of dense BNNs. The dilemma of how to cut down massive training costs remains, particularly given the requirement to learn about the uncertainty. To solve this challenge, we introduce Sparse Subspace Variational Inference (SSVI), the first fully sparse BNN framework that maintains a consistently highly sparse Bayesian model throughout the training and inference phases. Starting from a randomly initialized low-dimensional sparse subspace, our approach alternately optimizes the sparse subspace basis selection and its associated parameters. While basis selection is characterized as a non-differentiable problem, we approximate the optimal solution with a removal-and-addition strategy, guided by novel criteria based on weight distribution statistics. Our extensive experiments show that SSVI sets new benchmarks in crafting sparse BNNs, achieving, for instance, a 10-20x compression in model size with under 3% performance drop, and up to 20x FLOPs reduction during training compared with dense VI training. Remarkably, SSVI also demonstrates enhanced robustness to hyperparameters, reducing the need for intricate tuning in VI and occasionally even surpassing VI-trained dense BNNs on both accuracy and uncertainty metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Partially Stochastic Infinitely Deep Bayesian Neural NetworksSergio Calvo-Ordoñez, Matthieu Meunier, Francesco Piatti, Yuantao ShiICML 2024 · 7 citations
- BMRS: Bayesian Model Reduction for Structured PruningDustin Wright, Christian Igel, Raghavendra SelvanNeurIPS 2024 · 7 citations
- Coeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large ModelsZichen Miao, Wei Chen, Qiang QiuCVPR 2025
Builds on10
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 845 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen et al.ICLR 2020 · 292 citations
- Efficient and Scalable Bayesian Neural Nets with Rank-1 FactorsMichael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma et al.ICML 2020 · 239 citations
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 217 citations
Related papers
- Efficient Variational Inference for Sparse Deep Learning with Theoretical GuaranteeJincheng Bai, Qifan Song, Guang ChengNeurIPS 2020 · 55 citations
- SIKA-GP: Accelerating Gaussian Process Inference with Sparse Inducing Kernel Approximations for Bayesian Deep LearningWenyuan Zhao, Rui Tuo, Chao TianICML 2026
- Collapsed Inference for Bayesian Deep LearningZhe Zeng, Guy Van den BroeckNeurIPS 2023 · 10 citations
- Fast-BCNN: Massive Neuron Skipping in Bayesian Convolutional Neural NetworksQiyu Wan, Xin FuMICRO 2020 · 26 citations
- Shift-BNN: Highly-Efficient Probabilistic Bayesian Neural Network Training via Memory-Friendly Pattern RetrievingQiyu Wan, Haojun Xia, Xingyao Zhang, Lening Wang et al.MICRO 2021 · 9 citations
