Global Convergence and Rich Feature Learning in L-Layer Infinite-Width Neural Networks under μ Parametrization
Zixiang Chen, Greg Yang, Qingyue Zhao, Quanquan Gu
Abstract
Despite deep neural networks' powerful representation learning capabilities, theoretical understanding of how networks can simultaneously achieve meaningful feature learning and global convergence remains elusive. Existing approaches like the neural tangent kernel (NTK) are limited because features stay close to their initialization in this parametrization, leaving open questions about feature properties during substantial evolution. In this paper, we investigate the training dynamics of infinitely wide, L-layer neural networks using the tensor program (TP) framework. Specifically, we show that, when trained with stochastic gradient descent (SGD) under the Maximal Update parametrization (µP) and mild conditions on the activation function, SGD enables these networks to learn linearly independent features that substantially deviate from their initial values. This rich feature space captures relevant data information and ensures that any convergent point of the training process is a global minimum. Our analysis leverages both the interactions among features across layers and the properties of Gaussian random variables, providing new insights into deep representation learning. We further validate our theoretical findings through experiments on real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c98a55b8-ce4b-4cd2-b71d-6c2b111fe537Builds on13
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Infinite attention: NNGP and NTK for deep attention networksJiri Hron, Yasaman Bahri, Jascha Sohl-Dickstein, Roman NovakICML 2020 · 147 citations
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 140 citations
- A Generalized Neural Tangent Kernel Analysis for Two-layer Neural NetworksZixiang Chen, Yuan Cao, Quanquan Gu, Tong ZhangNeurIPS 2020 · 82 citations
Related papers
- Adaptive Optimization in the ∞-Width LimitEtai Littwin, Greg YangICLR 2023
- Tensor Programs IIb: Architectural Universality Of Neural Tangent Kernel Training DynamicsGreg Yang, Etai LittwinICML 2021 · 81 citations
- Real-Valued Backpropagation is Unsuitable for Complex-Valued Neural NetworksZhi-Hao Tan, Yi Xie, Yuan Jiang, Zhi-Hua ZhouNeurIPS 2022 · 16 citations
- Finite Depth and Width Corrections to the Neural Tangent KernelBoris Hanin, Mihai NicaICLR 2020 · 169 citations
- Neural Tangent Kernels Under Stochastic Data AugmentationJoshua DeOliveira, Sajal Chakroborty, Walter Gerych, Elke A. RundensteinerAAAI 2026
