Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural Networks
Ryo Karakida, Kazuki Osawa
摘要
Natural gradient descent (NGD) helps to accelerate the convergence of gradient descent dynamics, but it requires approximations in large-scale deep neural networks because of its high computational cost. Empirical studies have confirmed that some NGD methods with approximate Fisher information converge sufficiently fast in practice. Nevertheless, it remains unclear from the theoretical perspective why and under what conditions such heuristic approximations work well. In this work, we reveal that, under specific conditions, NGD with approximate Fisher information achieves the same fast convergence to global minima as exact NGD. We consider deep neural networks in the infinite-width limit, and analyze the asymptotic training dynamics of NGD in function space via the neural tangent kernel. In the function space, the training dynamics with the approximate Fisher information are identical to those with the exact Fisher information, and they converge quickly. The fast convergence holds in layer-wise approximations; for instance, in block diagonal approximation where each block corresponds to a layer as well as in block tri-diagonal and K-FAC approximations. We also find that a unit-wise approximation achieves the same fast convergence under some assumptions. All of these different approximations have an isotropic gradient in the function space, and this plays a fundamental role in achieving the same convergence properties in training. Thus, the current study gives a novel and unified theoretical foundation with which to understand NGD methods in deep learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- The Geometry of Neural Nets' Parameter Spaces Under ReparametrizationAgustinus Kristiadi, Felix Dangel, Philipp HennigNeurIPS 2023 · 被引用 21 次
- When does preconditioning help or hurt generalization?Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li 等ICLR 2021 · 被引用 11 次
- Rethinking Gauss-Newton for learning over-parameterized modelsMichael Arbel, Romain Menegaux, Pierre WolinskiNeurIPS 2023 · 被引用 10 次
- On the Parameterization of Second-Order Optimization Effective towards the Infinite WidthSatoki Ishikawa, Ryo KarakidaICLR 2024 · 被引用 10 次
- Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor GeneralizationDavide Buffelli, Jamie McGowan, Wangkun Xu, Alexandru Cioba 等NeurIPS 2024 · 被引用 6 次
相关 Paper
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 被引用 2 次
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
- HyLo: A Hybrid Low-Rank Natural Gradient Descent MethodBaorun Mu, Saeed Soori, Bugra Can, Mert Gürbüzbalaban 等SC 2022 · 被引用 3 次
- Rich Information is Affordable: A Systematic Performance Analysis of Second-order Optimization Using K-FACYuichiro Ueno, Kazuki Osawa, Yohei Tsuji, Akira Naruse 等KDD 2020 · 被引用 9 次
- Convolutional neural network training with distributed K-FACJ. Gregory Pauloski, Zhao Zhang, Lei Huang, Weijia Xu 等SC 2020 · 被引用 26 次
