Efficient Computation of Deep Nonlinear Infinite-Width Neural Networks that Learn Features
Greg Yang, Michael Santacroce, Edward J. Hu
Abstract
While a popular limit of infinite-width neural networks, the Neural Tangent Kernel (NTK) often exhibits performance gaps from finite-width neural networks on standard datasets, due to lack of feature learning. Although the feature learning maximal update limit, or μ-limit (Yang and Hu, 2020) of wide networks has closed the gap for 1-hidden-layer linear models, no one has been able to demonstrate this for deep nonlinear multi-layer perceptrons (MLP) because of μ-limit’s computational difficulty in this setting. Here, we solve this problem by proposing a novel feature learning limit, the π-limit, that bypasses the computational issues. The π-limit, in short, is the limit of a form of projected gradient descent, and the π-limit of an MLP is roughly another MLP where gradients are appended to weights during training. We prove its almost sure convergence with width using the Tensor Programs technique. We evaluate it on CIFAR10 and Omniglot against NTK as well as finite networks, finding the π-limit outperform finite-width models trained normally (without projection) in both settings, closing the performance gap between finite- and infinite-width neural networks previously left by NTK. Code for this work is available at github.com/santacml/pilim.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 201a91f6-1e41-4b72-9d0f-e072cfb4eebdCited by top-tier papers5
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 140 citations
- Width and Depth Limits Commute in Residual NetworksSoufiane Hayou, Greg YangICML 2023 · 23 citations
- Synergy and Symmetry in Deep Learning: Interactions between the Data, Model, and Inference AlgorithmLechao Xiao, Jeffrey PenningtonICML 2022 · 12 citations
- Convolutional Deep Kernel MachinesEdward Milsom, Ben Anson, Laurence AitchisonICLR 2024 · 6 citations
Related papers
- Adaptive Optimization in the ∞-Width LimitEtai Littwin, Greg YangICLR 2023
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Tensor Programs IIb: Architectural Universality Of Neural Tangent Kernel Training DynamicsGreg Yang, Etai LittwinICML 2021 · 81 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- Fast Finite Width Neural Tangent KernelRoman Novak, Jascha Sohl-Dickstein, Samuel S. SchoenholzICML 2022 · 72 citations
