Trade-Offs of Diagonal Fisher Information Matrix Estimators
Alexander Soen, Ke Sun
摘要
The Fisher information matrix can be used to characterize the local geometry of the parameter space of neural networks. It elucidates insightful theories and useful tools to understand and optimize neural networks. Given its high computational cost, practitioners often use random estimators and evaluate only the diagonal entries. We examine two popular estimators whose accuracy and sample complexity depend on their associated variances. We derive bounds of the variances and instantiate them in neural networks for regression and classification. We navigate trade-offs for both estimators based on analytical and numerical studies. We find that the variance quantities depend on the non-linearity wrt different parameter groups and should not be neglected when estimating the Fisher information.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Disentangling Consensus and Value-Specific Representations for Controllable Pluralistic Value Alignment of LLMsJianKui Zhou, Jing Yao, Xiaoyuan Yi, Peng Zhang 等ICML 2026
- Scalable Kronecker-Factored Fisher Approximation for Neural Network Parameter SensitivityViktoriia Chekalina, Daniil Moskovskiy, Tatyana Matveeva, Andrey Kuznetsov 等ICML 2026
- Deterministic Bounds and Random Estimates of Metric Tensors on NeuromanifoldsKe SunICLR 2026
- Domain Sensitive Federated Learning with Fisher-Informed PruningChenchen Lin, Wenhao Yuan, Zhengji Xu, Xuehe WangCVPR 2026
它引用的顶会 Paper8
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa 等AAAI 2021 · 被引用 358 次
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 被引用 217 次
- BackPACK: Packing more into BackpropFelix Dangel, Frederik Kunstner, Philipp HennigICLR 2020 · 被引用 114 次
- Understanding Approximate Fisher Information for Fast Convergence of Natural Gradient Descent in Wide Neural NetworksRyo Karakida, Kazuki OsawaNeurIPS 2020 · 被引用 39 次
相关 Paper
- On the Variance of the Fisher Information for Deep LearningAlexander Soen, Ke SunNeurIPS 2021 · 被引用 31 次
- Discriminating image representations with principal distortionsJenelle Feather, David Lipshutz, Sarah E. Harvey, Alex H. Williams 等ICLR 2025
- Information-Theoretic Local Minima Characterization and RegularizationZhiwei Jia, Hao SuICML 2020 · 被引用 22 次
- Gone Fishing: Neural Active Learning with Fisher EmbeddingsJordan T. Ash, Surbhi Goel, Akshay Krishnamurthy, Sham M. KakadeNeurIPS 2021 · 被引用 124 次
- Fishers for Free? Approximating the Fisher Information Matrix by Recycling the Squared Gradient AccumulatorYu Xin Li, Felix Dangel, Derek Tam, Colin RaffelICML 2025
