Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep Learning
Wu Lin, Valentin Duruisseaux, Melvin Leok, Frank Nielsen, Mohammad Emtiyaz Khan, Mark Schmidt
摘要
Riemannian submanifold optimization with momentum is computationally challenging because, to ensure that the iterates remain on the submanifold, we often need to solve difficult differential equations. Here, we simplify such difficulties for a class of sparse or structured symmetric positivedefinite matrices with the affine-invariant metric. We do so by proposing a generalized version of the Riemannian normal coordinates that dynamically orthonormalizes the metric and locally converts the problem into an unconstrained problem in the Euclidean space. We use our approach to simplify existing approaches for structured covariances and develop matrix-inverse-free 2 nd -order optimizers for deep learning with low precision by using only matrix multiplications. Because the set of SPD matrices forms a Riemannian manifold, one can use Riemannian gradient methods for SPD estimation, but this can be computationally infeasible in high-dimensions. This is because the methods often require full-rank matrix decomposition (see Table 1 ). Computations
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Kronecker-Factored Approximate Curvature for Modern Neural Network ArchitecturesRuna Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider 等NeurIPS 2023 · 被引用 62 次
- Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveWu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae 等ICML 2024 · 被引用 23 次
- Understanding and improving Shampoo and SOAP via Kullback-Leibler MinimizationWu Lin, Scott C. Lowe, Felix Dangel, Runa Eschenhagen 等ICLR 2026 · 被引用 15 次
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov 等ICML 2024 · 被引用 7 次
- SVRG and Beyond via Posterior CorrectionNico Daheim, Thomas Moellenhoff, James Ming Liang Ang, Mohammad Emtiyaz KhanICML 2026
它引用的顶会 Paper6
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- On Riemannian Optimization over Positive Definite Matrices with the Bures-Wasserstein GeometryAndi Han, Bamdev Mishra, Pratik Kumar Jawanpuria, Junbin GaoNeurIPS 2021 · 被引用 55 次
- Tractable structured natural-gradient descent using local parameterizationsWu Lin, Frank Nielsen, Mohammad Emtiyaz Khan, Mark SchmidtICML 2021 · 被引用 36 次
- Fast and Accurate Model ScalingPiotr Dollár, Mannat Singh, Ross B. GirshickCVPR 2021
- Rethinking Channel Dimensions for Efficient Model DesignDongyoon Han, Sangdoo Yun, Byeongho Heo, Youngjoon YooCVPR 2021
相关 Paper
- Riemannian coordinate descent algorithms on matrix manifoldsAndi Han, Pratik Jawanpuria, Bamdev MishraICML 2024 · 被引用 10 次
- Learning to Optimize on SPD ManifoldsZhi Gao, Yuwei Wu, Yunde Jia, Mehrtash HarandiCVPR 2020
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 被引用 139 次
- Learning a Gradient-free Riemannian Optimizer on Tangent SpacesXiaomeng Fan, Zhi Gao, Yuwei Wu, Yunde Jia 等AAAI 2021 · 被引用 8 次
- Decentralized Projected Riemannian Stochastic Recursive Momentum Method for Nonconvex OptimizationKangkang Deng, Jiang HuAAAI 2025 · 被引用 2 次
