Subquadratic Overparameterization for Shallow Neural Networks
Chaehwan Song, Ali Ramezani-Kebrya, Thomas Pethick, Armin Eftekhari, Volkan Cevher
摘要
Overparameterization refers to the important phenomenon where the width of a neural network is chosen such that learning algorithms can provably attain zero loss in nonconvex training. The existing theory establishes such global convergence using various initialization strategies, training modifications, and width scalings. In particular, the state-of-the-art results require the width to scale quadratically with the number of training data under standard initialization strategies used in practice for best generalization performance. In contrast, the most recent results obtain linear scaling either with requiring initializations that lead to the "lazy-training", or training only a single layer. In this work, we provide an analytical framework that allows us to adopt standard initialization strategies, possibly avoid lazy training, and train all layers simultaneously in basic shallow neural networks while attaining a desirable subquadratic scaling on the network width. We achieve the desiderata via Polyak-Łojasiewicz condition, smoothness, and standard assumptions on data, and use tools from random matrix theory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- BOME! Bilevel Optimization Made Easy: A Simple First-Order ApproachBo Liu, Mao Ye, Stephen Wright, Peter Stone 等NeurIPS 2022 · 被引用 170 次
- On Penalty Methods for Nonconvex Bilevel Optimization and First-Order Stochastic ApproximationJeongyeol Kwon, Dohyun Kwon, Stephen Wright, Robert D. NowakICLR 2024 · 被引用 61 次
- Memorization and Optimization in Deep Neural Networks with Minimum Over-parameterizationSimone Bombari, Mohammad Hossein Amani, Marco MondelliNeurIPS 2022 · 被引用 45 次
- Robustness in deep learning: The good (width), the bad (depth), and the ugly (initialization)Zhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2022 · 被引用 28 次
- On the Convergence of Encoder-only Shallow TransformersYongtao Wu, Fanghui Liu, Grigorios Chrysos, Volkan CevherNeurIPS 2023 · 被引用 17 次
它引用的顶会 Paper9
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 被引用 193 次
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 被引用 183 次
- Harnessing the Power of Infinitely Wide Deep Nets on Small-data TasksSanjeev Arora, Simon S. Du, Zhiyuan Li, Ruslan Salakhutdinov 等ICLR 2020 · 被引用 167 次
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu 等ICML 2020 · 被引用 85 次
相关 Paper
- Convergence Rates of Non-Convex Stochastic Gradient Descent Under a Generic Lojasiewicz Condition and Local SmoothnessKevin Scaman, Cédric Malherbe, Ludovic Dos SantosICML 2022 · 被引用 24 次
- Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal TopologyQuynh Nguyen, Marco MondelliNeurIPS 2020 · 被引用 82 次
- On global convergence of ResNets: From finite to infinite width using linear parameterizationRaphaël Barboni, Gabriel Peyré, François-Xavier VialardNeurIPS 2022 · 被引用 14 次
- Width Independent Bounds for the Local Lipschitz Constant of Deep Neural Networks at Random Initialization and after Lazy TrainingApostolos Evangelidis, Felix KrahmerICML 2026
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?Zixiang Chen, Yuan Cao, Difan Zou, Quanquan GuICLR 2021 · 被引用 29 次
