Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networks
Ziwei Ji, Matus Telgarsky
摘要
Recent theoretical work has guaranteed that overparameterized networks trained by gradient descent achieve arbitrarily low training error, and sometimes even low test error. The required width, however, is always polynomial in at least one of the sample size , the (inverse) target error , and the (inverse) failure probability . This work shows that iterations of gradient descent with training examples on two-layer ReLU networks of any width exceeding suffice to achieve a test misclassification error of . We also prove that stochastic gradient descent can achieve test error with polylogarithmic width and samples. The analysis relies upon the separation margin of the limiting kernel, which is guaranteed positive, can distinguish between true labels and random labels, and can give a tight sample-complexity analysis in the infinite-width setting
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper85
- Deep learning versus kernel learning: an empirical study of loss landscape geometry and the time evolution of the Neural Tangent KernelStanislav Fort, Gintare Karolina Dziugaite, Mansheej Paul, Sepideh Kharaghani 等NeurIPS 2020 · 被引用 255 次
- A Group-Theoretic Framework for Data AugmentationShuxiao Chen, Edgar Dobriban, Jane H. LeeNeurIPS 2020 · 被引用 254 次
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational LimitBoaz Barak, Benjamin L. Edelman, Surbhi Goel, Sham M. Kakade 等NeurIPS 2022 · 被引用 220 次
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 被引用 183 次
- High-dimensional Asymptotics of Feature Learning: How One Gradient Step Improves the RepresentationJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Zhichao Wang 等NeurIPS 2022 · 被引用 173 次
它引用的顶会 Paper1
相关 Paper
- How Much Over-parameterization Is Sufficient to Learn Deep ReLU Networks?Zixiang Chen, Yuan Cao, Difan Zou, Quanquan GuICLR 2021 · 被引用 29 次
- Feature selection and low test error in shallow low-rotation ReLU networksMatus TelgarskyICLR 2023
- Beyond NTK with Vanilla Gradient Descent: A Mean-Field Analysis of Neural Networks with Polynomial Width, Samples, and TimeArvind V. Mahankali, Haochen Zhang, Kefan Dong, Margalit Glasgow 等NeurIPS 2023 · 被引用 20 次
- On Convergence and Generalization of Dropout TrainingPoorya Mianjy, Raman AroraNeurIPS 2020 · 被引用 34 次
- A single gradient step finds adversarial examples on random two-layers neural networksSébastien Bubeck, Yeshwanth Cherapanamjeri, Gauthier Gidel, Remi Tachet des CombesNeurIPS 2021 · 被引用 31 次
