M-FAC: Efficient Matrix-Free Approximations of Second-Order Information
Elias Frantar, Eldar Kurtic, Dan Alistarh
摘要
Efficiently approximating local curvature information of the loss function is a key tool for optimization and compression of deep neural networks. Yet, most existing methods to approximate second-order information have high computational or storage costs, which can limit their practicality. In this work, we investigate matrix-free, linear-time approaches for estimating Inverse-Hessian Vector Products (IHVPs) for the case when the Hessian can be approximated as a sum of rank-one matrices, as in the classic approximation of the Hessian by the empirical Fisher matrix. We propose two new algorithms as part of a framework called M-FAC: the first algorithm is tailored towards network compression and can compute the IHVP for dimension , if the Hessian is given as a sum of rank-one matrices, using precomputation, cost for computing the IHVP, and query cost for any single element of the inverse Hessian. The second algorithm targets an optimization setting, where we wish to compute the product between the inverse Hessian, estimated over a sliding window of optimization steps, and a given gradient direction, as required for preconditioned SGD. We give an algorithm with cost for computing the IHVP and for adding or removing any gradient from the sliding window. These two algorithms yield state-of-the-art results for network pruning and optimization with lower computational overhead relative to existing second-order methods. Implementations are available at [9] and [17].
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningElias Frantar, Dan AlistarhNeurIPS 2022 · 被引用 440 次
- SpikingBERT: Distilling BERT to Train Spiking Language Models Using Implicit DifferentiationMalyaban Bal, Abhronil SenguptaAAAI 2024 · 被引用 78 次
- ZipLM: Inference-Aware Structured Pruning of Language ModelsEldar Kurtic, Elias Frantar, Dan AlistarhNeurIPS 2023 · 被引用 69 次
- Compressing LLMs: The Truth is Rarely Pure and Never SimpleAjay Kumar Jaiswal, Zhe Gan, Xianzhi Du, Bowen Zhang 等ICLR 2024 · 被引用 61 次
- The Emergence of Essential Sparsity in Large Pre-trained Models: The Weights that MatterAjay Jaiswal, Shiwei Liu, Tianlong Chen, Zhangyang WangNeurIPS 2023 · 被引用 57 次
它引用的顶会 Paper3
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa 等AAAI 2021 · 被引用 358 次
- Soft Threshold Weight Reparameterization for Learnable SparsityAditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman 等ICML 2020 · 被引用 266 次
- WoodFisher: Efficient Second-Order Approximation for Neural Network CompressionSidak Pal Singh, Dan AlistarhNeurIPS 2020 · 被引用 217 次
相关 Paper
- Error Feedback Can Accurately Compress PreconditionersIonut-Vlad Modoranu, Aleksei Kalinov, Eldar Kurtic, Elias Frantar 等ICML 2024 · 被引用 6 次
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
- THOR, Trace-based Hardware-driven Layer-Oriented Natural Gradient Descent ComputationMengyun Chen, Kai-Xin Gao, Xiaolei Liu, Zidong Wang 等AAAI 2021 · 被引用 7 次
- Rich Information is Affordable: A Systematic Performance Analysis of Second-order Optimization Using K-FACYuichiro Ueno, Kazuki Osawa, Yohei Tsuji, Akira Naruse 等KDD 2020 · 被引用 9 次
- Scalable Kronecker-Factored Fisher Approximation for Neural Network Parameter SensitivityViktoriia Chekalina, Daniil Moskovskiy, Tatyana Matveeva, Andrey Kuznetsov 等ICML 2026
