Lune

NeurIPS2020Top-tier venue

Global Convergence of Deep Networks with One Wide Layer Followed by Pyramidal Topology

Quynh Nguyen, Marco Mondelli

2020Year
82Citations
33Top-tier citations

Abstract

Recent works have shown that gradient descent can find a global minimum for over-parameterized neural networks where the widths of all the hidden layers scale polynomially with NN (NN being the number of training samples). In this paper, we prove that, for deep networks, a single layer of width NN following the input layer suffices to ensure a similar guarantee. In particular, all the remaining layers are allowed to have constant widths, and form a pyramidal topology. We show an application of our result to the widely used Xavier's initialization and obtain an over-parameterization requirement for the single wide layer of order N2.N^2.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Cited by top-tier papers33

Ask how each one uses it

Builds on3

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines