Variational Inference for Infinitely Deep Neural Networks
Achille Nazaret, David M. Blei
摘要
We introduce the unbounded depth neural network (UDN), an infinitely deep probabilistic model that adapts its complexity to the training data. The UDN contains an infinite sequence of hidden layers and places an unbounded prior on a truncation L, the layer from which it produces its data. Given a dataset of observations, the posterior UDN provides a conditional distribution of both the parameters of the infinite neural network and its truncation. We develop a novel variational inference algorithm to approximate this posterior, optimizing a distribution of the neural network weights and of the truncation depth L, and without any upper limit on L. To this end, the variational family has a special structure: it models neural network weights of arbitrary depth, and it dynamically creates or removes free variational parameters as its distribution of the truncation is optimized. (Unlike heuristic approaches to model search, it is solely through gradient-based optimization that this algorithm explores the space of truncations.) We study the UDN on real and synthetic data. We find that the UDN adapts its posterior depth to the dataset complexity; it outperforms standard neural networks of similar computational complexity; and it outperforms other approaches to infinite-depth neural networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Bayesian Adaptation of Network Depth and Width for Continual LearningJeevan Thapa, Rui LiICML 2024 · 被引用 7 次
- Adaptive Width Neural NetworksFederico Errica, Henrik Christiansen, Viktor Zaverkin, Mathias Niepert 等ICLR 2026 · 被引用 6 次
- Enhancing Diversity in Bayesian Deep Learning via Hyperspherical Energy Minimization of CKADavid Smerkous, Qinxun Bai, Fuxin LiNeurIPS 2024 · 被引用 3 次
- Adaptive Message Passing: A General Framework to Mitigate Oversmoothing, Oversquashing, and UnderreachingFederico Errica, Henrik Christiansen, Viktor Zaverkin, Takashi Maruyama 等ICML 2025
它引用的顶会 Paper4
- Multiscale Deep Equilibrium ModelsShaojie Bai, Vladlen Koltun, J. Zico KolterNeurIPS 2020 · 被引用 272 次
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 等ICLR 2020 · 被引用 254 次
- Depth Uncertainty in Neural NetworksJavier Antorán, James Urquhart Allingham, José Miguel Hernández-LobatoNeurIPS 2020 · 被引用 121 次
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process PerspectiveGeoff Pleiss, John P. CunninghamNeurIPS 2021 · 被引用 35 次
相关 Paper
- Partially Stochastic Infinitely Deep Bayesian Neural NetworksSergio Calvo-Ordoñez, Matthieu Meunier, Francesco Piatti, Yuantao ShiICML 2024 · 被引用 7 次
- Specifying Weight Priors in Bayesian Deep Neural Networks with Empirical BayesRanganath Krishnan, Mahesh Subedar, Omesh TickooAAAI 2020 · 被引用 65 次
- Masked Bayesian Neural Networks : Theoretical Guarantee and its Posterior InferenceInsung Kong, Dongyoon Yang, Jongjin Lee, Ilsang Ohn 等ICML 2023 · 被引用 8 次
- Liberty or Depth: Deep Bayesian Neural Nets Do Not Need Complex Weight Posterior ApproximationsSebastian Farquhar, Lewis Smith, Yarin GalNeurIPS 2020 · 被引用 47 次
- Effective Estimation of Deep Generative Language ModelsTom Pelsmaeker, Wilker AzizACL 2020 · 被引用 5 次
