Adaptive Optimization in the ∞-Width Limit
Etai Littwin, Greg Yang
Abstract
Recent works have developed detailed understanding of large neural networks' behaviors via their infinite-width limits, e.g., the neural tangent kernel (NTK) and the feature learning () limits. These theories were developed for stochastic gradient descent. Yet, in practice, all large NN are trained using Adam or other adaptive gradient optimizers (AGO), which are not covered by such previous works. Here, we close this gap via the Tensor Programs framework. Specifically, for deep MLPs, we derive the NTK and parametrizations as well as their infinite-width limits. We find 1) The NTK limit of AGO, in contrast to that of SGD, now depends nonlinearly on the loss derivative but nevertheless still fails to learn features; 2) this is fixed by the limit of AGO (as in the case of SGD). To obtain these results, we extend the Tensor Programs language with a new instruction that allows one to express the gradient processing done by AGOs.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 670358db-413b-47af-93c4-90783ab0e6d5Cited by top-tier papers2
- Preconditioning Neural Tangent Kernel for Adaptive OptimizationXiyuan Yang, Wenxuan Bao, Katherine Tieu, Jingrui HeICML 2026
- Dichotomy of Feature Learning and Unlearning: Fast-Slow Analysis on Neural Networks with Stochastic Gradient DescentShota Imai, Sota Nishiyama, Masaaki ImaizumiICML 2026
Related papers
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Efficient Computation of Deep Nonlinear Infinite-Width Neural Networks that Learn FeaturesGreg Yang, Michael Santacroce, Edward J. HuICLR 2022 · 9 citations
- Global Convergence and Rich Feature Learning in L-Layer Infinite-Width Neural Networks under μ ParametrizationZixiang Chen, Greg Yang, Qingyue Zhao, Quanquan GuICML 2025
- Tensor Programs IIb: Architectural Universality Of Neural Tangent Kernel Training DynamicsGreg Yang, Etai LittwinICML 2021 · 81 citations
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 140 citations
