An Improved Empirical Fisher Approximation for Natural Gradient Descent
Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip C. Woodland
Abstract
Approximate Natural Gradient Descent (NGD) methods are an important family of optimisers for deep learning models, which use approximate Fisher information matrices to pre-condition gradients during training. The empirical Fisher (EF) method approximates the Fisher information matrix empirically by reusing the per-sample gradients collected during back-propagation. Despite its ease of implementation, the EF approximation has its theoretical and practical limitations. This paper investigates the inversely-scaled projection issue of EF, which is shown to be a major cause of its poor empirical approximation quality. An improved empirical Fisher (iEF) method is proposed to address this issue, which is motivated as a generalised NGD method from a loss reduction perspective, meanwhile retaining the practical convenience of EF. The exact iEF and EF methods are experimentally evaluated using practical deep learning setups, including widely-used setups for parameter-efficient fine-tuning of pre-trained models (T5-base with LoRA and Prompt-Tuning on GLUE tasks, and ViT with LoRA for CIFAR100). Optimisation experiments show that applying exact iEF directly as an optimiser provides strong convergence and generalisation. It achieves the best test performance and the lowest training loss for the majority of the tasks, even when compared to well-tuned AdamW/Adafactor baselines. Additionally, under a novel empirical evaluation framework, the proposed iEF method shows consistently better approximation quality to exact Natural Gradient updates than both the EF and the more expensive sampled Fisher methods, meanwhile demonstrating the superior property of being robust to the choice of damping across tasks and training stages. Improving existing approximate NGD optimisers with iEF is expected to lead to better convergence and robustness. Furthermore, the iEF method also serves as a better approximation method to the Fisher information matrix itself, which enables the improvement of a variety of Fisher-based methods, not limited to the scope of optimisation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Can Fine-Tuning Erase Edits? On the Fragile Coexistence of Knowledge Editing and Fine-tuningYinjie Cheng, Paul Youssef, Christin Seifert, Jörg Schlötterer et al.KDD 2026 · 2 citations
- Decomposing the Basic Abilities of Large Language Models: Mitigating Cross-Task Interference in Multi-Task Instruct-TuningBing Wang, Ximing Li, Changchun Li, Jinjin Chi et al.ICML 2026 · 1 citation
- Disentangling Consensus and Value-Specific Representations for Controllable Pluralistic Value Alignment of LLMsJianKui Zhou, Jing Yao, Xiaoyuan Yi, Peng Zhang et al.ICML 2026
- Scalable Kronecker-Factored Fisher Approximation for Neural Network Parameter SensitivityViktoriia Chekalina, Daniil Moskovskiy, Tatyana Matveeva, Andrey Kuznetsov et al.ICML 2026
- Randomized Advantage Transformation (RAT): Computing Natural Policy Gradients via Direct BackpropagationMingfei SunICML 2026
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Robust Speech Recognition via Large-Scale Weak SupervisionAlec Radford, Jong Wook Kim, Tao Xu, Greg Brockman et al.ICML 2023 · 6,966 citations
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 217 citations
Related papers
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 2 citations
- Rich Information is Affordable: A Systematic Performance Analysis of Second-order Optimization Using K-FACYuichiro Ueno, Kazuki Osawa, Yohei Tsuji, Akira Naruse et al.KDD 2020 · 9 citations
- AdaFisher: Adaptive Second Order Optimization via Fisher InformationDamien Martins Gomes, Yanlei Zhang, Eugene Belilovsky, Guy Wolf et al.ICLR 2025
- Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural NetworksRyo Karakida, Kazuki OsawaNeurIPS 2020 · 39 citations
- LoRA-DA: Data-Aware Initialization for Low-Rank Adaptation via Asymptotic AnalysisQingyue Zhang, Chang Chu, Tianren Peng, Qi Li et al.ICML 2026
