On the Parameterization of Second-Order Optimization Effective towards the Infinite Width
Satoki Ishikawa, Ryo Karakida
摘要
Second-order optimization has been developed to accelerate the training of deep neural networks and it is being applied to increasingly larger-scale models. In this study, towards training on further larger scales, we identify a specific parameterization for second-order optimization that promotes feature learning in a stable manner even if the network width increases significantly. Inspired by a maximal update parameterization, we consider a one-step update of the gradient and reveal the appropriate scales of hyperparameters including random initialization, learning rates, and damping terms. Our approach covers two major second-order optimization algorithms, K-FAC and Shampoo, and we demonstrate that our parameterization achieves higher generalization performance in feature learning. In particular, it enables us to transfer the hyperparameters across models with different widths.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Scaling Exponents Across Parameterizations and OptimizersKatie E. Everett, Lechao Xiao, Mitchell Wortsman, Alexander A. Alemi 等ICML 2024 · 被引用 59 次
- μPC: Scaling Predictive Coding to 100+ Layer NetworksFrancesco Innocenti, El Mehdi Achour, Christopher L. BuckleyNeurIPS 2025 · 被引用 19 次
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei 等NeurIPS 2025 · 被引用 17 次
- μP2: Effective Sharpness Aware Minimization Requires Layerwise Perturbation ScalingMoritz Haas, Jin Xu, Volkan Cevher, Leena Chennuru VankadaraNeurIPS 2024 · 被引用 12 次
- FedMuon: Federated Learning with Bias-corrected LMO-based OptimizationYuki Takezawa, Anastasia Koloskova, Xiaowen Jiang, Sebastian U. StichICLR 2026 · 被引用 9 次
它引用的顶会 Paper17
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- Knowledge distillation: A good teacher is patient and consistentLucas Beyer, Xiaohua Zhai, Amélie Royer, Larisa Markeeva 等CVPR 2022 · 被引用 215 次
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- Small-scale proxies for large-scale Transformer training instabilitiesMitchell Wortsman, Peter J. Liu, Lechao Xiao, Katie E. Everett 等ICLR 2024 · 被引用 162 次
- Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural NetworksBlake Bordelon, Cengiz PehlevanNeurIPS 2022 · 被引用 140 次
相关 Paper
- Eva: Practical Second-order Optimization with Kronecker-vectorized ApproximationLin Zhang, Shaohuai Shi, Bo LiICLR 2023
- Tensor Normal Training for Deep Learning ModelsYi Ren, Donald GoldfarbNeurIPS 2021 · 被引用 36 次
- Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveWu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae 等ICML 2024 · 被引用 23 次
- Gradient Descent on Neurons and its Link to Approximate Second-order OptimizationFrederik BenzingICML 2022 · 被引用 31 次
- Studying K-FAC Heuristics by Viewing Adam through a Second-Order LensRoss M. Clarke, José Miguel Hernández-LobatoICML 2024 · 被引用 2 次
