Trade-Offs of Diagonal Fisher Information Matrix Estimators
Alexander Soen, Ke Sun
Abstract
The Fisher information matrix can be used to characterize the local geometry of the parameter space of neural networks. It elucidates insightful theories and useful tools to understand and optimize neural networks. Given its high computational cost, practitioners often use random estimators and evaluate only the diagonal entries. We examine two popular estimators whose accuracy and sample complexity depend on their associated variances. We derive bounds of the variances and instantiate them in neural networks for regression and classification. We navigate trade-offs for both estimators based on analytical and numerical studies. We find that the variance quantities depend on the non-linearity wrt different parameter groups and should not be neglected when estimating the Fisher information.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c0f560f-1bcd-4769-bfb2-f4ee39b16d1fCited by top-tier papers4
- Disentangling Consensus and Value-Specific Representations for Controllable Pluralistic Value Alignment of LLMsJianKui Zhou, Jing Yao, Xiaoyuan Yi, Peng Zhang et al.ICML 2026
- Scalable Kronecker-Factored Fisher Approximation for Neural Network Parameter SensitivityViktoriia Chekalina, Daniil Moskovskiy, Tatyana Matveeva, Andrey Kuznetsov et al.ICML 2026
- Deterministic Bounds and Random Estimates of Metric Tensors on NeuromanifoldsKe SunICLR 2026
- Domain Sensitive Federated Learning with Fisher-Informed PruningChenchen Lin, Wenhao Yuan, Zhengji Xu, Xuehe WangCVPR 2026
Builds on8
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa et al.AAAI 2021 · 358 citations
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 217 citations
- BackPACK: Packing more into BackpropFelix Dangel, Frederik Kunstner, Philipp HennigICLR 2020 · 114 citations
- Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural NetworksRyo Karakida, Kazuki OsawaNeurIPS 2020 · 39 citations
Related papers
- On the Variance of the Fisher Information for Deep LearningAlexander Soen, Ke SunNeurIPS 2021 · 31 citations
- Discriminating image representations with principal distortionsJenelle Feather, David Lipshutz, Sarah E. Harvey, Alex H. Williams et al.ICLR 2025
- Information-Theoretic Local Minima Characterization and RegularizationZhiwei Jia, Hao SuICML 2020 · 22 citations
- Gone Fishing: Neural Active Learning with Fisher EmbeddingsJordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, Sham M. KakadeNeurIPS 2021 · 124 citations
- Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient AccumulatorYu Xin Li, Felix Dangel, Derek Tam, Colin RaffelICML 2025
