The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich Regimes
Alexander B. Atanasov, Blake Bordelon, Sabarish Sainathan, Cengiz Pehlevan
摘要
For small training set sizes , the generalization error of wide neural networks is well-approximated by the error of an infinite width neural network (NN), either in the kernel or mean-field/feature-learning regime. However, after a critical sample size , we empirically find the finite-width network generalization becomes worse than that of the infinite width network. In this work, we empirically study the transition from infinite-width behavior to this variance limited regime as a function of sample size and network width . We find that finite-size effects can become relevant for very small dataset sizes on the order of for polynomial regression with ReLU networks. We discuss the source of these effects using an argument based on the variance of the NN's final neural tangent kernel (NTK). This transition can be pushed to larger by enhancing feature learning or by ensemble averaging the networks. We find that the learning curve for regression with the final NTK is an accurate approximation of the NN learning curve. Using this, we provide a toy model which also exhibits scaling and has -dependent benefits from feature learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- TRAK: Attributing Model Behavior at ScaleSung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc 等ICML 2023 · 被引用 260 次
- Resolving Discrepancies in Compute-Optimal Scaling of Language ModelsTomer Porian, Mitchell Wortsman, Jenia Jitsev, Ludwig Schmidt 等NeurIPS 2024 · 被引用 94 次
- Grokking as the transition from lazy to rich training dynamicsTanishq Kumar, Blake Bordelon, Samuel J. Gershman, Cengiz PehlevanICLR 2024 · 被引用 86 次
- A Dynamical Model of Neural Scaling LawsBlake Bordelon, Alexander B. Atanasov, Cengiz PehlevanICML 2024 · 被引用 84 次
- Dynamics of Finite Width Kernel and Prediction Fluctuations in Mean Field Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2023 · 被引用 56 次
相关 Paper
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 被引用 169 次
- On the linearity of large non-linear models: when and why the tangent kernel is constantChaoyue Liu, Libin Zhu, Mikhail BelkinNeurIPS 2020 · 被引用 183 次
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 被引用 167 次
- What can linearized neural networks actually say about generalization?Guillermo Ortiz-Jiménez, Seyed-Mohsen Moosavi-Dezfooli, Pascal FrossardNeurIPS 2021 · 被引用 62 次
- On the Generalization Power of Overfitted Two-Layer Neural Tangent Kernel ModelsPeizhong Ju, Xiaojun Lin, Ness B. ShroffICML 2021 · 被引用 13 次
