Training Bayesian Neural Networks with Sparse Subspace Variational Inference
Junbo Li, Zichen Miao, Qiang Qiu, Ruqi Zhang
摘要
Bayesian neural networks (BNNs) offer uncertainty quantification but come with the downside of substantially increased training and inference costs. Sparse BNNs have been investigated for efficient inference, typically by either slowly introducing sparsity throughout the training or by post-training compression of dense BNNs. The dilemma of how to cut down massive training costs remains, particularly given the requirement to learn about the uncertainty. To solve this challenge, we introduce Sparse Subspace Variational Inference (SSVI), the first fully sparse BNN framework that maintains a consistently highly sparse Bayesian model throughout the training and inference phases. Starting from a randomly initialized low-dimensional sparse subspace, our approach alternately optimizes the sparse subspace basis selection and its associated parameters. While basis selection is characterized as a non-differentiable problem, we approximate the optimal solution with a removal-and-addition strategy, guided by novel criteria based on weight distribution statistics. Our extensive experiments show that SSVI sets new benchmarks in crafting sparse BNNs, achieving, for instance, a 10-20x compression in model size with under 3% performance drop, and up to 20x FLOPs reduction during training compared with dense VI training. Remarkably, SSVI also demonstrates enhanced robustness to hyperparameters, reducing the need for intricate tuning in VI and occasionally even surpassing VI-trained dense BNNs on both accuracy and uncertainty metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Partially Stochastic Infinitely Deep Bayesian Neural NetworksSergio Calvo-Ordoñez, Matthieu Meunier, Francesco Piatti, Yuantao ShiICML 2024 · 被引用 7 次
- BMRS: Bayesian Model Reduction for Structured PruningDustin Wright, Christian Igel, Raghavendra SelvanNeurIPS 2024 · 被引用 7 次
- Coeff-Tuning: A Graph Filter Subspace View for Tuning Attention-Based Large ModelsZichen Miao, Wei Chen, Qiang QiuCVPR 2025
它引用的顶会 Paper10
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Cyclical Stochastic Gradient MCMC for Bayesian Deep LearningRuqi Zhang, Chunyuan Li, Jianyi Zhang, Changyou Chen 等ICLR 2020 · 被引用 292 次
- Efficient and Scalable Bayesian Neural Nets with Rank-1 FactorsMichael Dusenberry, Ghassen Jerfel, Yeming Wen, Yi-An Ma 等ICML 2020 · 被引用 239 次
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 被引用 217 次
相关 Paper
- Efficient Variational Inference for Sparse Deep Learning with Theoretical GuaranteeJincheng Bai, Qifan Song, Guang ChengNeurIPS 2020 · 被引用 55 次
- SIKA-GP: Accelerating Gaussian Process Inference with Sparse Inducing Kernel Approximations for Bayesian Deep LearningWenyuan Zhao, Rui Tuo, Chao TianICML 2026
- Collapsed Inference for Bayesian Deep LearningZhe Zeng, Guy Van den BroeckNeurIPS 2023 · 被引用 10 次
- Fast-BCNN: Massive Neuron Skipping in Bayesian Convolutional Neural NetworksQiyu Wan, Xin FuMICRO 2020 · 被引用 26 次
- Shift-BNN: Highly-Efficient Probabilistic Bayesian Neural Network Training via Memory-Friendly Pattern RetrievingQiyu Wan, Haojun Xia, Xingyao Zhang, Lening Wang 等MICRO 2021 · 被引用 9 次
