Lune

ICLR2025顶会

How Gradient descent balances features: A dynamical analysis for two-layer neural networks

Zhenyu Zhu, Fanghui Liu, Volkan Cevher

出版方
2025年份
1顶会引用

摘要

This paper investigates the fundamental regression task of learning k neurons (a.k.a. teachers) from Gaussian input, using two-layer ReLU neural networks with width m (a.k.a. students) and m, k = O(1), trained via gradient descent under proper initialization and a small step-size. Our analysis follows a threephase structure: alignment after weak recovery, tangential growth, and local convergence, providing deeper insights into the learning dynamics of gradient descent (GD). We prove the global convergence at the rate of O(T -3 ) for the zero loss of excess risk. Additionally, our results show that GD automatically groups and balances student neurons, revealing an implicit bias toward achieving the minimum "balanced" ℓ 2 -norm in the solution. Our work extends beyond previous studies in exact-parameterization setting (m = k = 1, (Yehudai and Ohad, 2020)) and single-neuron setting (m ≥ k = 1, (Xu and Du, 2023)). The key technical challenge lies in handling the interactions between multiple teachers and students during training, which we address by refining the alignment analysis in Phase 1 and introducing a new dynamic system analysis for tangential components in Phase 2. Our results pave the way for further research on optimizing neural network training dynamics and understanding implicit biases in more complex architectures.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖