Lune

ICLR2024顶会

Early Neuron Alignment in Two-layer ReLU Networks with Small Initialization

Hancheng Min, Enrique Mallada, René Vidal

2024年份
31被引次数
18顶会引用

摘要

This paper studies the problem of training a two-layer ReLU network for binary classification using gradient flow with small initialization. We consider a training dataset with well-separated input vectors: Any pair of input data with the same label are positively correlated, and any pair with different labels are negatively correlated. Our analysis shows that, during the early phase of training, neurons in the first layer try to align with either the positive data or the negative data, depending on its corresponding weight on the second layer. A careful analysis of the neurons' directional dynamics allows us to provide an O(log⁡nμ)\mathcal{O}(\frac{\log n}{\sqrt{\mu}}) upper bound on the time it takes for all neurons to achieve good alignment with the input data, where nn is the number of data points and μ\mu measures how well the data are separated. After the early alignment phase, the loss converges to zero at a O(1t)\mathcal{O}(\frac{1}{t}) rate, and the weight matrix on the first layer is approximately low-rank. Numerical experiments on the MNIST dataset illustrate our theoretical findings.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 8d8d4613-c272-48b1-9f03-3aef2abb6daa

引用它的顶会 Paper18

问问它们各自怎么用它

它引用的顶会 Paper17

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖