Are wider nets better given the same number of parameters?
Anna Golubeva, Guy Gur-Ari, Behnam Neyshabur
摘要
Empirical studies demonstrate that the performance of neural networks improves with increasing number of parameters. In most of these studies, the number of parameters is increased by increasing the network width. This begs the question: Is the observed improvement due to the larger number of parameters, or is it due to the larger width itself? We compare different ways of increasing model width while keeping the number of parameters constant. We show that for models initialized with a random, static sparsity pattern in the weight tensors, network width is the determining factor for good performance, while the number of weights is secondary, as long as the model achieves high training accuarcy. As a step towards understanding this effect, we analyze these models in the framework of Gaussian Process kernels. We find that the distance between the sparse finite-width model kernel and the infinite-width kernel at initialization is indicative of model performance. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Gradient Flow in Sparse Neural Networks and How Lottery Tickets WinUtku Evci, Yani Ioannou, Cem Keskin, Yann N. DauphinAAAI 2022 · 被引用 106 次
- Stochastic Training is Not Necessary for GeneralizationJonas Geiping, Micah Goldblum, Phillip Pope, Michael Moeller 等ICLR 2022 · 被引用 83 次
- Why Random Pruning Is All We Need to Start SparseAdvait Harshal Gadhikar, Sohom Mukherjee, Rebekka BurkholzICML 2023 · 被引用 33 次
- PHEW : Constructing Sparse Networks that Learn Fast and Generalize Well without Training DataShreyas Malakarjun Patil, Constantine DovrolisICML 2021 · 被引用 26 次
- Masks, Signs, And Learning Rate RewindingAdvait Harshal Gadhikar, Rebekka BurkholzICLR 2024 · 被引用 15 次
它引用的顶会 Paper5
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Pruning Neural Networks at Initialization: Why Are We Missing the Mark?Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICLR 2021 · 被引用 261 次
- Fast Sparse ConvNetsErich Elsen, Marat Dukhan, Trevor Gale, Karen SimonyanCVPR 2020
相关 Paper
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process PerspectiveGeoff Pleiss, John P. CunninghamNeurIPS 2021 · 被引用 35 次
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- More is Better: when Infinite Overparameterization is Optimal and Overfitting is ObligatoryJames B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail BelkinICLR 2024 · 被引用 7 次
- Unveiling The Matthew Effect Across Channels: Assessing Layer Width Sufficiency via Weight Norm VarianceYiting Chen, Jiazi Bu, Junchi YanNeurIPS 2024 · 被引用 4 次
- Training-Free Determination of Network Width via Neural Tangent KernelTatsumi Sunada, Toshihiko Yamasaki, Atsuto MakiICLR 2026
