Lune

ICML2026顶会

Singular Bayesian Neural Networks

Mame Diarra Toure, David A Stephens

2026年份
2被引次数
1顶会引用

摘要

Bayesian neural networks promise calibrated uncertainty but require O(mn)O(mn) parameters for standard mean-field Gaussian posteriors. We argue this cost is often unnecessary, particularly when weight matrices exhibit fast singular value decay. By parameterizing weights as W=AB⊤W = AB^{\top} with A∈Rm×rA \in \mathbb{R}^{m \times r}, B∈Rn×rB \in \mathbb{R}^{n \times r}, we induce a posterior that is singular with respect to the Lebesgue measure, concentrating on the rank-rr manifold. This singularity captures structured weight correlations through shared latent factors, geometrically distinct from mean-field's independence assumption. We derive PAC-Bayes generalization bounds whose complexity term scales as r(m+n)\sqrt{r(m+n)} instead of mn\sqrt{m n}, and prove loss bounds that decompose the error into optimization and rank-induced bias using the Eckart-Young-Mirsky theorem. We further adapt recent Gaussian complexity bounds for low-rank deterministic networks to Bayesian predictive means. Empirically, across MLPs, LSTMs, and Transformers on standard benchmarks, our method achieves competitive predictive performance while using up to 33×33\times fewer parameters than 5-member Deep Ensembles. It substantially improves OOD detection and often improves calibration relative to mean-field and perturbation baselines, while Deep Ensembles can still be stronger on in-distribution likelihood-based metrics.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper9

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖