Tractable structured natural-gradient descent using local parameterizations
Wu Lin, Frank Nielsen, Mohammad Emtiyaz Khan, Mark Schmidt
Abstract
Natural-gradient descent (NGD) on structured parameter spaces (e.g., low-rank covariances) is computationally challenging due to difficult Fisher-matrix computations. We address this issue by using local-parameter coordinates to obtain a flexible and efficient NGD method that works well for a wide-variety of structured parameterizations. We show four applications where our method (1) generalizes the exponential natural evolutionary strategy, (2) recovers existing Newton-like algorithms, (3) yields new structured second-order algorithms via matrix groups, and (4) gives new algorithms to learn covariances of Gaussian and Wishart-based distributions. We show results on a range of problems from deep learning, variational inference, and evolution strategies. Our work opens a new direction for scalable structured geometric methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 358b6dcf-281f-4c25-b093-38db25d33faaCited by top-tier papers8
- Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveWu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae et al.ICML 2024 · 23 citations
- The Geometry of Neural Nets' Parameter Spaces Under ReparametrizationAgustinus Kristiadi, Felix Dangel, Philipp HennigNeurIPS 2023 · 21 citations
- Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep LearningWu Lin, Valentin Duruisseaux, Melvin Leok, Frank Nielsen et al.ICML 2023 · 13 citations
- Understanding Stochastic Natural Gradient Variational InferenceKaiwen Wu, Jacob R. GardnerICML 2024 · 11 citations
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov et al.ICML 2024 · 7 citations
Builds on1
Related papers
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 2 citations
- Riemannian coordinate descent algorithms on matrix manifoldsAndi Han, Pratik Jawanpuria, Bamdev MishraICML 2024 · 10 citations
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li et al.CVPR 2021
- Fisher-Legendre (FishLeg) optimization of deep neural networksJezabel R. Garcia, Federica Freddi, Stathi Fotiadis, Maolin Li et al.ICLR 2023
- The Quotient Bayesian Learning RuleMykola Lukashchuk, Raphaël Trésor, Wouter W. L. Nuijten, Ismail Senöz et al.NeurIPS 2025 · 1 citation
