Beyond IID weights: sparse and low-rank deep Neural Networks are also Gaussian Processes
Thiziri Nait Saada, Alireza Naderi, Jared Tanner
摘要
The infinitely wide neural network has been proven a useful and manageable mathematical model that enables the understanding of many phenomena appearing in deep learning. One example is the convergence of random deep networks to Gaussian processes that allows a rigorous analysis of the way the choice of activation function and network weights impacts the training dynamics. In this paper, we extend the seminal proof of Matthews et al. (2018) to a larger class of initial weight distributions (which we call PSEUDO-IID), including the established cases of IID and orthogonal weights, as well as the emerging low-rank and structured sparse settings celebrated for their computational speed-up benefits. We show that fully-connected and convolutional networks initialized with PSEUDO-IID distributions are all effectively equivalent up to their variance. Using our results, one can identify the Edge-of-Chaos for a broader class of neural networks and tune them at criticality in order to enhance their training. Moreover, they enable the posterior distribution of Bayesian Neural Networks to be tractable across these various initialization schemes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Compacter: Efficient Low-Rank Hypercomplex Adapter LayersRabeeh Karimi Mahabadi, James Henderson, Sebastian RuderNeurIPS 2021 · 被引用 700 次
- Neural Collapse Under MSE Loss: Proximity to and Dynamics on the Central PathX. Y. Han, Vardan Papyan, David L. DonohoICLR 2022 · 被引用 182 次
- Monarch: Expressive Structured Matrices for Efficient and Accurate TrainingTri Dao, Beidi Chen, Nimit Sharad Sohoni, Arjun D. Desai 等ICML 2022 · 被引用 125 次
相关 Paper
- Width and Depth Limits Commute in Residual NetworksSoufiane Hayou, Greg YangICML 2023 · 被引用 23 次
- Deep Kernel Posterior Learning under Infinite Variance Prior WeightsJorge Loría, Anindya BhadraICLR 2025
- Large-width functional asymptotics for deep Gaussian neural networksDaniele Bracale, Stefano Favaro, Sandra Fortini, Stefano PeluchettiICLR 2021 · 被引用 2 次
- Critical feature learning in deep neural networksKirsten Fischer, Javed Lindner, David Dahmen, Zohar Ringel 等ICML 2024 · 被引用 15 次
- Precise characterization of the prior predictive distribution of deep ReLU networksLorenzo Noci, Gregor Bachmann, Kevin Roth, Sebastian Nowozin 等NeurIPS 2021 · 被引用 36 次
