Lune

NeurIPS2021顶会

Explicit loss asymptotics in the gradient descent training of neural networks

Maksim Velikanov, Dmitry Yarotsky

出版方
2021年份
19被引次数
9顶会引用

摘要

Current theoretical results on optimization trajectories of neural networks trained by gradient descent typically have the form of rigorous but potentially loose bounds on the loss values. In the present work we take a different approach and show that the learning trajectory of a wide network in a lazy training regime can be characterized by an explicit asymptotic at large training times. Specifically, the leading term in the asymptotic expansion of the loss behaves as a power law L(t) ∼ Ct -ξ with exponent ξ expressed only through the data dimension, the smoothness of the activation function, and the class of function being approximated. Our results are based on spectral analysis of the integral operator representing the linearized evolution of a large network trained on the expected loss. Importantly, the techniques we employ do not require a specific form of the data distribution, for example Gaussian, thus making our findings sufficiently universal. 10 -1 10 1 10 3 10 5 t 10 -6 10 -5 10 -4 10 -3 10 -2 10 -1 10 0 Loss GP. d=2 GP. d=4 Ind. d=2 Ind. d=4 10 0 10 1 10 2 10 3 10 4 n 10 -8 10 -6 10 -4 10 -2 10 0 Eigenvalues λn d=2 d=4 10 0 10 1 10 2 10 3 10 4 n 10 -7 10 -5 10 -3 10 -1 Coefficient partial sums sn GP. d=2 GP. d=4 Ind. d=2 Ind. d=4

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper9

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖