Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep Learning
Wu Lin, Valentin Duruisseaux, Melvin Leok, Frank Nielsen, Mohammad Emtiyaz Khan, Mark Schmidt
Abstract
Riemannian submanifold optimization with momentum is computationally challenging because, to ensure that the iterates remain on the submanifold, we often need to solve difficult differential equations. Here, we simplify such difficulties for a class of sparse or structured symmetric positivedefinite matrices with the affine-invariant metric. We do so by proposing a generalized version of the Riemannian normal coordinates that dynamically orthonormalizes the metric and locally converts the problem into an unconstrained problem in the Euclidean space. We use our approach to simplify existing approaches for structured covariances and develop matrix-inverse-free 2 nd -order optimizers for deep learning with low precision by using only matrix multiplications. Because the set of SPD matrices forms a Riemannian manifold, one can use Riemannian gradient methods for SPD estimation, but this can be computationally infeasible in high-dimensions. This is because the methods often require full-rank matrix decomposition (see Table 1 ). Computations
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e591bd84-23de-417b-be3f-dec9ebd40bffCited by top-tier papers5
- Kronecker-Factored Approximate Curvature for Modern Neural Network ArchitecturesRuna Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider et al.NeurIPS 2023 · 62 citations
- Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveWu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae et al.ICML 2024 · 23 citations
- Understanding and improving Shampoo and SOAP via Kullback-Leibler MinimizationWu Lin, Scott C. Lowe, Felix Dangel, Runa Eschenhagen et al.ICLR 2026 · 15 citations
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov et al.ICML 2024 · 7 citations
- SVRG and Beyond via Posterior CorrectionNico Daheim, Thomas Moellenhoff, James Ming Liang Ang, Mohammad Emtiyaz KhanICML 2026
Builds on6
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- On Riemannian Optimization over Positive Definite Matrices with the Bures-Wasserstein GeometryAndi Han, Bamdev Mishra, Pratik Kumar Jawanpuria, Junbin GaoNeurIPS 2021 · 55 citations
- Tractable structured natural-gradient descent using local parameterizationsWu Lin, Frank Nielsen, Mohammad Emtiyaz Khan, Mark SchmidtICML 2021 · 36 citations
- Fast and Accurate Model ScalingPiotr Dollár, Mannat Singh, Ross B. GirshickCVPR 2021
- Rethinking Channel Dimensions for Efficient Model DesignDongyoon Han, Sangdoo Yun, Byeongho Heo, Youngjoon YooCVPR 2021
Related papers
- Riemannian coordinate descent algorithms on matrix manifoldsAndi Han, Pratik Jawanpuria, Bamdev MishraICML 2024 · 10 citations
- Learning to Optimize on SPD ManifoldsZhi Gao, Yuwei Wu, Yunde Jia, Mehrtash HarandiCVPR 2020
- Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley TransformJun Li, Fuxin Li, Sinisa TodorovicICLR 2020 · 139 citations
- Learning a Gradient-free Riemannian Optimizer on Tangent SpacesXiaomeng Fan, Zhi Gao, Yuwei Wu, Yunde Jia et al.AAAI 2021 · 8 citations
- Decentralized Projected Riemannian Stochastic Recursive Momentum Method for Nonconvex OptimizationKangkang Deng, Jiang HuAAAI 2025 · 2 citations
