More is Better: when Infinite Overparameterization is Optimal and Overfitting is Obligatory
James B. Simon, Dhruva Karkada, Nikhil Ghosh, Mikhail Belkin
摘要
In our era of enormous neural networks, empirical progress has been driven by the philosophy that more is better. Recent deep learning practice has found repeatedly that larger model size, more data, and more computation (resulting in lower training loss) improves performance. In this paper, we give theoretical backing to these empirical observations by showing that these three properties hold in random feature (RF) regression, a class of models equivalent to shallow networks with only the last layer trained.
Concretely, we first show that the test risk of RF regression decreases monotonically with both the number of features and the number of samples, provided the ridge penalty is tuned optimally. In particular, this implies that infinite width RF architectures are preferable to those of any finite width. We then proceed to demonstrate that, for a large class of tasks characterized by powerlaw eigenstructure, training to near-zero training loss is obligatory: near-optimal performance can only be achieved when the training error is much smaller than the test error. Grounding our theory in real-world data, we find empirically that standard computer vision tasks with convolutional neural tangent kernels clearly fall into this class. Taken together, our results tell a simple, testable story of the benefits of overparameterization, overfitting, and more data in random feature models.
How might such theoretical results look? Consider the well-tested observation that wider networks virtually always achieve better performance, so long as they are properly tuned (Kaplan et al., 2020;Hoffmann et al., 2022;Yang et al., 2022). Let E te (n, w, θ) denote the expected test error of a network with width w and training hyperparameters θ when trained on n samples from an arbitrary distribution. A satisfactory explanation for this observation might be a hypothetical theorem which states the following:
If
Such a result would do much to bring deep learning theory up to date with practice. In this work, we take a first step towards this general result by proving it in the special case of RF regression -that is, for shallow networks with only the second layer trained. Our Theorem 1 states that, for RF regression, more features (as well as more data) is better, and thus infinite width is best. To our knowledge, this is the first analysis directly showing that for arbitrary tasks, wider is better for networks of a certain architecture.
How might a comparable result for overfitting look? It is by now established wisdom that optimal performance in many domains is achieved when training deep networks to nearly the point of
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Entity Insertion in Multilingual Linked Corpora: The Case of WikipediaTomás Feith, Akhil Arora, Martin Gerlach, Debjit Paul 等EMNLP 2024 · 被引用 1 次
- The φ Curve: The Shape of Generalization through the Lens of Norm-based Capacity ControlYichen Wang, Yudong Chen, Lorenzo Rosasco, Fanghui LiuNeurIPS 2025
它引用的顶会 Paper22
- A Universal Law of Robustness via IsoperimetrySébastien Bubeck, Mark SellkeNeurIPS 2021 · 被引用 260 次
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee 等ICLR 2020 · 被引用 254 次
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 被引用 245 次
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
相关 Paper
- Theoretical Limitations of Ensembles in the Age of OverparameterizationNiclas Dern, John Patrick Cunningham, Geoff PleissICML 2025
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 被引用 84 次
- Implicit Regularization of Random Feature ModelsArthur Jacot, Berfin Simsek, Francesco Spadaro, Clément Hongler 等ICML 2020 · 被引用 83 次
- Bayes-optimal Learning of Deep Random Networks of Extensive-widthHugo Cui, Florent Krzakala, Lenka ZdeborováICML 2023 · 被引用 49 次
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 被引用 133 次
