Lune

ICML2022Top-tier venue

On Non-local Convergence Analysis of Deep Linear Networks

Kun Chen, Dachao Lin, Zhihua Zhang

2022Year
1Citations

Abstract

In this paper, we follow Eftekhari [12] 's work to give a non-local convergence analysis of deep linear networks. Specifically, we consider optimizing deep linear networks which have a layer with one neuron under quadratic loss. We describe the convergent point of trajectories with arbitrary starting point under gradient flow, including the paths which converge to one of the saddle points or the original point. We also show specific convergence rates of trajectories that converge to the global minimizer by stages. To achieve these results, this paper mainly extends the machinery in [12] to provably identify the rank-stable set and the global minimizer convergent set. We also give specific examples to show the necessity of our definitions. Crucially, as far as we know, our results appear to be the first to give a non-local global analysis of linear neural networks from arbitrary initialized points, rather than the lazy training regime which has dominated the literature of neural networks, and restricted benign initialization in [12] . We also note that extending our results to general linear networks without one hidden neuron assumption remains a challenging open problem. * Equal Contribution.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 442e6dfa-fef9-4b88-b167-4a51a91deffd

Builds on6

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines