Lune

NeurIPS2021Top-tier venue

Explicit loss asymptotics in the gradient descent training of neural networks

Maksim Velikanov, Dmitry Yarotsky

2021Year
19Citations
9Top-tier citations

Abstract

Current theoretical results on optimization trajectories of neural networks trained by gradient descent typically have the form of rigorous but potentially loose bounds on the loss values. In the present work we take a different approach and show that the learning trajectory of a wide network in a lazy training regime can be characterized by an explicit asymptotic at large training times. Specifically, the leading term in the asymptotic expansion of the loss behaves as a power law L(t) ∼ Ct -ξ with exponent ξ expressed only through the data dimension, the smoothness of the activation function, and the class of function being approximated. Our results are based on spectral analysis of the integral operator representing the linearized evolution of a large network trained on the expected loss. Importantly, the techniques we employ do not require a specific form of the data distribution, for example Gaussian, thus making our findings sufficiently universal. 10 -1 10 1 10 3 10 5 t 10 -6 10 -5 10 -4 10 -3 10 -2 10 -1 10 0 Loss GP. d=2 GP. d=4 Ind. d=2 Ind. d=4 10 0 10 1 10 2 10 3 10 4 n 10 -8 10 -6 10 -4 10 -2 10 0 Eigenvalues λn d=2 d=4 10 0 10 1 10 2 10 3 10 4 n 10 -7 10 -5 10 -3 10 -1 Coefficient partial sums sn GP. d=2 GP. d=4 Ind. d=2 Ind. d=4

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 913f3387-94d5-4dfa-8ff6-e89e71687a8a

Cited by top-tier papers9

Ask how each one uses it

Builds on8

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines