Lune

ICLR2020顶会

Bounds on Over-Parameterization for Guaranteed Existence of Descent Paths in Shallow ReLU Networks

Arsalan Sharif-Nassab, Saber Salehkaleybar, S. Jamaloddin Golestani

出版方
2020年份
12被引次数
3顶会引用

摘要

We study the landscape of squared loss in neural networks with one-hidden layer and ReLU activation functions. Let mm and dd be the widths of hidden and input layers, respectively. We show that there exist poor local minima with positive curvature for some training sets of size n≥m+2d−2n\geq m+2d-2. By positive curvature of a local minimum, we mean that within a small neighborhood the loss function is strictly increasing in all directions. Consequently, for such training sets, there are initialization of weights from which there is no descent path to global optima. It is known that for n≤mn\le m, there always exist descent paths to global optima from all initial weights. In this perspective, our results provide a somewhat sharp characterization of the over-parameterization required for "existence of descent paths" in the loss landscape.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper3

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖