Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective
Wu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae, Richard E. Turner, Alireza Makhzani
摘要
Adaptive gradient optimizers like Adam(W) are the default training algorithms for many deep learning architectures, such as transformers. Their diagonal preconditioner is based on the gradient outer product which is incorporated into the parameter update via a square root. While these methods are often motivated as approximate second-order methods, the square root represents a fundamental difference. In this work, we investigate how the behavior of adaptive methods changes when we remove the root, i.e., strengthen their second-order motivation. Surprisingly, we find that such square-root-free adaptive methods close the generalization gap to SGD on convolutional architectures, while maintaining their root-based counterpart's performance on transformers. The second-order perspective also has practical benefits for developing non-diagonal methods that can incorporate arbitrary curvature approximations through the concept of preconditioner invariance. In contrast to root-based methods like Shampoo, root-free counterparts work well and fast with half-precision since they do not require numerically unstable matrix root decompositions and inversions. Overall, our findings provide new insights into the development of adaptive methods and raise important questions regarding the overlooked role of adaptivity in their success. (experiment code: https://github.com/yorkerlin/remove-the-square-root optimizer code: https://github.com/f-dangel/sirfshampoo)
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren 等NeurIPS 2025 · 被引用 58 次
- COSMOS: A Hybrid Adaptive Optimizer for Efficient Training of Large Language ModelsLiming Liu, Zhenghao Xu, Zixuan Zhang, Hao Kang 等ICLR 2026 · 被引用 27 次
- Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its PreconditionerRuna Eschenhagen, Aaron Defazio, Tsung-Hsien Lee, Richard E. Turner 等NeurIPS 2025 · 被引用 24 次
- Bayesian Online Natural Gradient (BONG)Matt Jones, Peter G. Chang, Kevin P. MurphyNeurIPS 2024 · 被引用 20 次
- Understanding and improving Shampoo and SOAP via Kullback-Leibler MinimizationWu Lin, Scott C. Lowe, Felix Dangel, Runa Eschenhagen 等ICLR 2026 · 被引用 15 次
它引用的顶会 Paper13
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- Focal Modulation NetworksJianwei Yang, Chunyuan Li, Xiyang Dai, Jianfeng GaoNeurIPS 2022 · 被引用 494 次
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim 等NeurIPS 2020 · 被引用 397 次
相关 Paper
- Combining Axes Preconditioners through Kronecker Approximation for Deep LearningSai Surya Duvvuri, Devvrit, Rohan Anil, Cho-Jui Hsieh 等ICLR 2024 · 被引用 16 次
- SOAP: Improving and Stabilizing Shampoo using Adam for Language ModelingNikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira 等ICLR 2025
- 4-bit Shampoo for Memory-Efficient Network TrainingSike Wang, Pan Zhou, Jia Li, Hua HuangNeurIPS 2024 · 被引用 19 次
- A New Perspective on Shampoo's PreconditionerDepen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach 等ICLR 2025
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan 等ICML 2026 · 被引用 1 次
