Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient Accumulator
Yu Xin Li, Felix Dangel, Derek Tam, Colin Raffel
Abstract
The diagonal of a model's Fisher Information Matrix (the "Fisher diagonal") has frequently been used as a way to measure parameter sensitivity. Typically, the Fisher diagonal is estimated via squared sampled gradients of the model's likelihood with respect to its parameters, averaged over a few hundred or thousand examples -a process which incurs nontrivial computational costs. At the same time, adaptive gradient methods like the ubiquitous Adam optimizer compute a moving average of the squared gradient over the course of training. This paper therefore explores whether an approximation of the Fisher diagonal can be obtained "for free" by recycling the squared gradient accumulator that has already been computed over the course of training. Through a comprehensive set of experiments covering five applications of the Fisher diagonal, we demonstrate that the "Squisher" (Squared gradient accumulator as an approximation of the Fisher) consistently performs similarly to the Fisher diagonal while outperforming baseline methods. Additionally, we clarify the exact differences between the Squisher and the Fisher diagonal and provide empirical quantification of their respective impact.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2144d1b-2458-4209-9be4-b414e7bedf88Cited by top-tier papers4
- Dataless Weight Disentanglement in Task Arithmetic via Kronecker-Factored Approximate CurvatureAngelo Porrello, Pietro Buzzega, Felix Dangel, Thomas Sommariva et al.ICLR 2026 · 6 citations
- The Unseen Frontier: Pushing the Limits of LLM Sparsity with Surrogate-Free ADMMKwanhee Lee, Hyeondo Jang, Dongyeop Lee, Dan Alistarh et al.ICLR 2026 · 5 citations
- The Appeal and Reality of Recycling LoRAs with Adaptive MergingHaokun Liu, Gyung Hyun Je, Marco Ciccone, Zhenlin Xu et al.ICML 2026 · 1 citation
- Understanding the Learning Phases in Self-Supervised Learning via Critical PeriodsJanghyeon Lee, Philipe A. Dias, Yao-Yi Chiang, Dalton D. LungaICLR 2026
Builds on18
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- AdaFisher: Adaptive Second Order Optimization via Fisher InformationDamien Martins Gomes, Yanlei Zhang, Eugene Belilovsky, Guy Wolf et al.ICLR 2025
- Rich Information is Affordable: A Systematic Performance Analysis of Second-order Optimization Using K-FACYuichiro Ueno, Kazuki Osawa, Yohei Tsuji, Akira Naruse et al.KDD 2020 · 9 citations
- IO-Adam: Rethinking Memory-Efficient Adaptive Optimizers from Gradient ComputationYiting Chen, Zongwei Huo, Junchi YanICML 2026
- A New Perspective on Shampoo's PreconditionerDepen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach et al.ICLR 2025
- Trade-Offs of Diagonal Fisher Information Matrix EstimatorsAlexander Soen, Ke SunNeurIPS 2024
