Tractable structured natural-gradient descent using local parameterizations
Wu Lin, Frank Nielsen, Mohammad Emtiyaz Khan, Mark Schmidt
摘要
Natural-gradient descent (NGD) on structured parameter spaces (e.g., low-rank covariances) is computationally challenging due to difficult Fisher-matrix computations. We address this issue by using local-parameter coordinates to obtain a flexible and efficient NGD method that works well for a wide-variety of structured parameterizations. We show four applications where our method (1) generalizes the exponential natural evolutionary strategy, (2) recovers existing Newton-like algorithms, (3) yields new structured second-order algorithms via matrix groups, and (4) gives new algorithms to learn covariances of Gaussian and Wishart-based distributions. We show results on a range of problems from deep learning, variational inference, and evolution strategies. Our work opens a new direction for scalable structured geometric methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order PerspectiveWu Lin, Felix Dangel, Runa Eschenhagen, Juhan Bae 等ICML 2024 · 被引用 23 次
- The Geometry of Neural Nets' Parameter Spaces Under ReparametrizationAgustinus Kristiadi, Felix Dangel, Philipp HennigNeurIPS 2023 · 被引用 21 次
- Simplifying Momentum-based Positive-definite Submanifold Optimization with Applications to Deep LearningWu Lin, Valentin Duruisseaux, Melvin Leok, Frank Nielsen 等ICML 2023 · 被引用 13 次
- Understanding Stochastic Natural Gradient Variational InferenceKaiwen Wu, Jacob R. GardnerICML 2024 · 被引用 11 次
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov 等ICML 2024 · 被引用 7 次
它引用的顶会 Paper1
相关 Paper
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 被引用 2 次
- Riemannian coordinate descent algorithms on matrix manifoldsAndi Han, Pratik Jawanpuria, Bamdev MishraICML 2024 · 被引用 10 次
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
- Fisher-Legendre (FishLeg) optimization of deep neural networksJezabel R. Garcia, Federica Freddi, Stathi Fotiadis, Maolin Li 等ICLR 2023
- The Quotient Bayesian Learning RuleMykola Lukashchuk, Raphaël Trésor, Wouter W. L. Nuijten, Ismail Senöz 等NeurIPS 2025 · 被引用 1 次
