Variational Deep Learning via Implicit Regularization
Jonathan Wenger, Beau Coker, Juraj Marusic, John Patrick Cunningham
摘要
Modern deep learning models generalize remarkably well in-distribution, despite being overparametrized and trained with little to no explicit regularization. Instead, current theory credits implicit regularization imposed by the choice of architecture, hyperparameters, and optimization procedure. However, deep neural networks can be surprisingly non-robust, resulting in overconfident predictions and poor out-of-distribution generalization. Bayesian deep learning addresses this via model averaging, but typically requires significant computational resources as well as carefully elicited priors to avoid overriding the benefits of implicit regularization. Instead, in this work, we propose to regularize variational neural networks solely by relying on the implicit bias of (stochastic) gradient descent. We theoretically characterize this inductive bias in overparametrized linear models as generalized variational inference and demonstrate the importance of the choice of parametrization. Empirically, our approach demonstrates strong in-and outof-distribution performance without additional hyperparameter tuning and with minimal computational overhead. Bayesian Deep Learning Approximate Bayesian techniques like the Laplace approximation [24] [25] [26] , stochastic weight averaging [27, 28] , deep ensembles [29], and variational approaches [30] [31] [32] [33] attempt to address the aforementioned shortcomings of deep learning by learning a distribution over functions as opposed to merely a point estimate. The idea being that a weighted combination of models, all of which achieve low training error, generalizes more robustly while at the same time providing uncertainty quantification. Variational Inference In Bayesian inference this weighted combination is defined by the posterior distribution p(w | X, y) ∝ p(y | X, w)p(w) over weights, induced by a likelihood p(y | w) and
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper24
- Bayesian Deep Learning and a Probabilistic Perspective of GeneralizationAndrew Gordon Wilson, Pavel IzmailovNeurIPS 2020 · 被引用 845 次
- Simple and Principled Uncertainty Estimation with Deterministic Deep Learning via Distance AwarenessJeremiah Z. Liu, Zi Lin, Shreyas Padhy, Dustin Tran 等NeurIPS 2020 · 被引用 604 次
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 被引用 458 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
相关 Paper
- Structured Dropout Variational Inference for Bayesian Neural NetworksSon Nguyen, Duong Nguyen, Khai Nguyen, Khoat Than 等NeurIPS 2021 · 被引用 11 次
- Convergence Rates of Variational Inference in Sparse Deep LearningBadr-Eddine Chérief-AbdellatifICML 2020 · 被引用 43 次
- Walsh-Hadamard Variational Inference for Bayesian Deep LearningSimone Rossi, Sébastien Marmin, Maurizio FilipponeNeurIPS 2020 · 被引用 17 次
- Deep Variational Implicit ProcessesLuis A. Ortega, Simón Rodríguez Santana, Daniel Hernández-LobatoICLR 2023 · 被引用 15 次
- Implicit Neural Representation Inference for Low-Dimensional Bayesian Deep LearningPanagiotis Dimitrakopoulos, Giorgos Sfikas, Christophoros NikouICLR 2024
