Lune

ICLR2024顶会

Implicit bias of SGD in L2-regularized linear DNNs: One-way jumps from high to low rank

Zihan Wang, Arthur Jacot

2024年份
27被引次数
18顶会引用

摘要

The L2L_{2}-regularized loss of Deep Linear Networks (DLNs) with more than one hidden layers has multiple local minima, corresponding to matrices with different ranks. In tasks such as matrix completion, the goal is to converge to the local minimum with the smallest rank that still fits the training data. While rank-underestimating minima can be avoided since they do not fit the data, GD might get stuck at rank-overestimating minima. We show that with SGD, there is always a probability to jump from a higher rank minimum to a lower rank one, but the probability of jumping back is zero. More precisely, we define a sequence of sets B1⊂B2⊂⋯⊂BRB_{1}\subset B_{2}\subset\cdots\subset B_{R} so that BrB_{r} contains all minima of rank rr or less (and not more) that are absorbing for small enough ridge parameters λ\lambda and learning rates η\eta: SGD has prob. 0 of leaving BrB_{r}, and from any starting point there is a non-zero prob. for SGD to go in BrB_{r}.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext aa89b396-977d-40e6-93dc-e939a2ebcfa5

引用它的顶会 Paper18

问问它们各自怎么用它

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖