On the Parameterization of Second-Order Optimization Effective towards the Infinite Width
Satoki Ishikawa, Ryo Karakida
Abstract
Second-order optimization has been developed to accelerate the training of deep neural networks and it is being applied to increasingly larger-scale models. In this study, towards training on further larger scales, we identify a specific parameterization for second-order optimization that promotes feature learning in a stable manner even if the network width increases significantly. Inspired by a maximal update parameterization, we consider a one-step update of the gradient and reveal the appropriate scales of hyperparameters including random initialization, learning rates, and damping terms. Our approach covers two major second-order optimization algorithms, K-FAC and Shampoo, and we demonstrate that our parameterization achieves higher generalization performance in feature learning. In particular, it enables us to transfer the hyperparameters across models with different widths.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Scaling Exponents Across Parameterizations and OptimizersKatie E. Everett, Lechao Xiao, Mitchell Wortsman, Alexander A. Alemi et al.ICML 2024 · 59 citations
- μPC: Scaling Predictive Coding to 100+ Layer NetworksFrancesco Innocenti, El Mehdi Achour, Christopher L. BuckleyNeurIPS 2025 · 19 citations
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei et al.NeurIPS 2025 · 17 citations
- μP2: Effective Sharpness Aware Minimization Requires Layerwise Perturbation ScalingMoritz Haas, Jin Xu, Volkan Cevher, Leena Chennuru VankadaraNeurIPS 2024 · 12 citations
- FedMuon: Federated Learning with Bias-corrected LMO-based OptimizationYuki Takezawa, Anastasia Koloskova, Xiaowen Jiang, Sebastian U. StichICLR 2026 · 9 citations
Builds on17
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 242 citations
- Knowledge distillation: A good teacher is patient and consistentLucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva et al.CVPR 2022 · 215 citations
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Small-scale proxies for large-scale Transformer training instabilitiesMitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett et al.ICLR 2024 · 162 citations
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 140 citations
Related papers
- Eva: Practical Second-order Optimization with Kronecker-vectorized ApproximationLin Zhang, Shaohuai Shi, Bo LiICLR 2023
- Tensor Normal Training for Deep Learning ModelsYi Ren, Donald GoldfarbNeurIPS 2021 · 36 citations
- Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveWu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae et al.ICML 2024 · 23 citations
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 31 citations
- Studying K-FAC Heuristics by Viewing Adam through a Second-Order LensRoss M. Clarke, José Miguel Hernández-LobatoICML 2024 · 2 citations
