Lune

ICLR2020Top-tier venue

Bounds on Over-Parameterization for Guaranteed Existence of Descent Paths in Shallow ReLU Networks

Arsalan Sharif-Nassab, Saber Salehkaleybar, S. Jamaloddin Golestani

2020Year
12Citations
3Top-tier citations

Abstract

We study the landscape of squared loss in neural networks with one-hidden layer and ReLU activation functions. Let mm and dd be the widths of hidden and input layers, respectively. We show that there exist poor local minima with positive curvature for some training sets of size n≥m+2d−2n\geq m+2d-2. By positive curvature of a local minimum, we mean that within a small neighborhood the loss function is strictly increasing in all directions. Consequently, for such training sets, there are initialization of weights from which there is no descent path to global optima. It is known that for n≤mn\le m, there always exist descent paths to global optima from all initial weights. In this perspective, our results provide a somewhat sharp characterization of the over-parameterization required for "existence of descent paths" in the loss landscape.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get d711faf8-b102-4e8d-83c9-3aaea34a2e26

Cited by top-tier papers3

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines