Understanding and improving Shampoo and SOAP via Kullback-Leibler Minimization
Wu Lin, Scott C. Lowe, Felix Dangel, Runa Eschenhagen, Zikun Xu, Roger B. Grosse
摘要
Shampoo and its efficient variant, SOAP, employ structured second-moment estimations and have shown strong performance for training neural networks (NNs). In practice, however, Shampoo typically requires step-size grafting with Adam to be competitive, and SOAP mitigates this by applying Adam in Shampoo's eigenbasis -- at the cost of additional memory overhead from Adam in both methods. Prior analyses have largely relied on the Frobenius norm to motivate these estimation schemes. We instead recast their estimation procedures as covariance estimation under Kullback-Leibler (KL) divergence minimization, revealing a previously overlooked theoretical limitation and motivating principled redesigns. Building on this perspective, we develop and , practical schemes that match or exceed the performance of Shampoo and SOAP in NN pre-training while achieving SOAP-level per-iteration runtime. Notably, KL-Shampoo does not rely on Adam to attain competitive performance, eliminating the memory overhead introduced by Adam. Across our experiments, KL-Shampoo consistently outperforms SOAP, Shampoo, and even KL-SOAP, establishing the KL-based approach as a promising foundation for designing structured methods in NN optimization. An implementation of KL-Shampoo/KL-SOAP is available at https://github.com/yorkerlin/KL-Methods
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its PreconditionerRuna Eschenhagen, Aaron Defazio, Tsung-Hsien Lee, Richard E. Turner 等NeurIPS 2025 · 被引用 24 次
- Adam Reduces a Unique Form of Sharpness: Theoretical Insights Near the Minimizer ManifoldXinghan Li, Haodong Wen, Kaifeng LyuNeurIPS 2025 · 被引用 6 次
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan 等ICML 2026 · 被引用 1 次
- FOAM: Frequency and Operator-Error Based Adaptive Damping Method for Reducing Staleness-Oriented Error for ShampooKyunghun Nam, Sumyeong AhnICML 2026
- Exploiting weight-space symmetries for approximating curvatureArtem Artemev, Rui Xia, Benjamin M. Boyd, Youjing Yu 等ICML 2026
它引用的顶会 Paper9
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren 等NeurIPS 2025 · 被引用 58 次
- Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its PreconditionerRuna Eschenhagen, Aaron Defazio, Tsung-Hsien Lee, Richard E. Turner 等NeurIPS 2025 · 被引用 24 次
- Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveWu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae 等ICML 2024 · 被引用 23 次
- Combining Axes Preconditioners through Kronecker Approximation for Deep LearningSai Surya Duvvuri, Devvrit, Rohan Anil, Cho-Jui Hsieh 等ICLR 2024 · 被引用 16 次
- Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep LearningWu Lin, Valentin Duruisseaux, Melvin Leok, Frank Nielsen 等ICML 2023 · 被引用 13 次
相关 Paper
- SOAP: Improving and Stabilizing Shampoo using Adam for Language ModelingNikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira 等ICLR 2025
- A New Perspective on Shampoo's PreconditionerDepen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach 等ICLR 2025
- Memory-Efficient 4-bit Preconditioned Stochastic OptimizationJingyang Li, Kuangyu Ding, Kim-Chuan Toh, Pan ZhouICCV 2025 · 被引用 1 次
- Sketchy: Memory-efficient Adaptive Regularization with Frequent DirectionsVladimir Feinberg, Xinyi Chen, Y. Jennifer Sun, Rohan Anil 等NeurIPS 2023 · 被引用 21 次
- Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided PreconditioningHuan Li, Yiming Dong, Zhouchen LinICML 2026
