Lune

ICLR2023Top-tier venue

The Onset of Variance-Limited Behavior for Networks in the Lazy and Rich Regimes

Alexander B. Atanasov, Blake Bordelon, Sabarish Sainathan, Cengiz Pehlevan

2023Year
4Citations
21Top-tier citations

Abstract

For small training set sizes PP, the generalization error of wide neural networks is well-approximated by the error of an infinite width neural network (NN), either in the kernel or mean-field/feature-learning regime. However, after a critical sample size P∗P^*, we empirically find the finite-width network generalization becomes worse than that of the infinite width network. In this work, we empirically study the transition from infinite-width behavior to this variance limited regime as a function of sample size PP and network width NN. We find that finite-size effects can become relevant for very small dataset sizes on the order of P∗∼NP^* \sim \sqrt{N} for polynomial regression with ReLU networks. We discuss the source of these effects using an argument based on the variance of the NN's final neural tangent kernel (NTK). This transition can be pushed to larger PP by enhancing feature learning or by ensemble averaging the networks. We find that the learning curve for regression with the final NTK is an accurate approximation of the NN learning curve. Using this, we provide a toy model which also exhibits P∗∼NP^* \sim \sqrt{N} scaling and has PP-dependent benefits from feature learning.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext d471389b-0aa0-4714-af0e-4a0741b78da1

Cited by top-tier papers21

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines