Are wider nets better given the same number of parameters?
Anna Golubeva, Guy Gur-Ari, Behnam Neyshabur
Abstract
Empirical studies demonstrate that the performance of neural networks improves with increasing number of parameters. In most of these studies, the number of parameters is increased by increasing the network width. This begs the question: Is the observed improvement due to the larger number of parameters, or is it due to the larger width itself? We compare different ways of increasing model width while keeping the number of parameters constant. We show that for models initialized with a random, static sparsity pattern in the weight tensors, network width is the determining factor for good performance, while the number of weights is secondary, as long as the model achieves high training accuarcy. As a step towards understanding this effect, we analyze these models in the framework of Gaussian Process kernels. We find that the distance between the sparse finite-width model kernel and the infinite-width kernel at initialization is indicative of model performance. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2e07fd0-4c78-434f-b532-554aa8df199fCited by top-tier papers15
- Gradient Flow in Sparse Neural Networks and How Lottery Tickets WinUtku Evci, Yani Ioannou, Cem Keskin, Yann N. DauphinAAAI 2022 · 106 citations
- Stochastic Training is Not Necessary for GeneralizationJonas Geiping, Micah Goldblum, Phillip Pope, Michael Moeller et al.ICLR 2022 · 83 citations
- Why Random Pruning Is All We Need to Start SparseAdvait Harshal Gadhikar, Sohom Mukherjee, Rebekka BurkholzICML 2023 · 33 citations
- PHEW : Constructing Sparse Networks that Learn Fast and Generalize Well without Training DataShreyas Malakarjun Patil, Constantine DovrolisICML 2021 · 26 citations
- Masks, Signs, And Learning Rate RewindingAdvait Harshal Gadhikar, Rebekka BurkholzICLR 2024 · 15 citations
Builds on5
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Pruning Neural Networks at Initialization: Why Are We Missing the Mark?Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICLR 2021 · 261 citations
- Fast Sparse ConvNetsErich Elsen, Marat Dukhan, Trevor Gale, Karen SimonyanCVPR 2020
Related papers
- The Limitations of Large Width in Neural Networks: A Deep Gaussian Process PerspectiveGeoff Pleiss, John P. CunninghamNeurIPS 2021 · 35 citations
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- More is Better: when Infinite Overparameterization is Optimal and Overfitting is ObligatoryJames B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail BelkinICLR 2024 · 7 citations
- Unveiling The Matthew Effect Across Channels: Assessing Layer Width Sufficiency via Weight Norm VarianceYiting Chen, Jiazi Bu, Junchi YanNeurIPS 2024 · 4 citations
- Training-Free Determination of Network Width via Neural Tangent KernelTatsumi Sunada, Toshihiko Yamasaki, Atsuto MakiICLR 2026
