AdaFisher: Adaptive Second Order Optimization via Fisher Information
Damien Martins Gomes, Yanlei Zhang, Eugene Belilovsky, Guy Wolf, Mahdi S. Hosseini
摘要
First-order optimization methods are currently the mainstream in training deep neural networks (DNNs). Optimizers like Adam incorporate limited curvature information by employing the diagonal matrix preconditioning of the stochastic gradient during the training. Despite their widespread, second-order optimization algorithms exhibit superior convergence properties compared to their first-order counterparts e.g. Adam and SGD. However, their practicality in training DNNs is still limited due to increased per-iteration computations compared to the first-order methods. We present AdaFisher-an adaptive second-order optimizer that leverages a diagonal block-Kronecker approximation of the Fisher information matrix for adaptive gradient preconditioning. AdaFisher aims to bridge the gap between enhanced convergence/generalization capabilities and computational efficiency in second-order optimization framework for training DNNs. Despite the slow pace of second-order optimizers, we showcase that AdaFisher can be reliably adopted for image classification, language modeling and stands out for its stability and robustness in hyper-parameter tuning. We demonstrate that AdaFisher outperforms the SOTA optimizers in terms of both accuracy and convergence speed. Code is available from https://github.com/AtlasAnalyticsLab/AdaFisher .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- SPICE: Submodular Penalized Information-Conflict Selection for Efficient Large Language Model TrainingPowei Chang, Jinpeng Zhang, Bowen Chen, Chenyu Wang 等ICLR 2026 · 被引用 5 次
- KOALA++: Efficient Kalman-Based Optimization with Gradient-Covariance ProductsZixuan Xia, Aram Davtyan, Paolo FavaroNeurIPS 2025 · 被引用 2 次
- Dynamic Momentum Recalibration in Online Gradient LearningZhipeng Yao, Rui Yu, Guisong Chang, Ying Li 等CVPR 2026 · 被引用 1 次
- Enhancing Optimizer Stability: Momentum Adaptation of The NGN Step-sizeRustem Islamov, Niccolò Ajroldi, Antonio Orvieto, Aurélien LucchiNeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper26
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
相关 Paper
- A Layer-Wise Natural Gradient Optimizer for Training Deep Neural NetworksXiaolei Liu, Shaoshuai Li, Kaixin Gao, Binfeng WangNeurIPS 2024 · 被引用 2 次
- Rich Information is Affordable: A Systematic Performance Analysis of Second-order Optimization Using K-FACYuichiro Ueno, Kazuki Osawa, Yohei Tsuji, Akira Naruse 等KDD 2020 · 被引用 9 次
- A Trace-restricted Kronecker-Factored Approximation to Natural GradientKai-Xin Gao, Xiao-Lei Liu, Zheng-Hai Huang, Min Wang 等AAAI 2021 · 被引用 13 次
- SKFAC: Training Neural Networks With Faster Kronecker-Factored Approximate CurvatureZedong Tang, Fenlong Jiang, Maoguo Gong, Hao Li 等CVPR 2021
- THOR, Trace-based Hardware-driven Layer-Oriented Natural Gradient Descent ComputationMengyun Chen, Kai-Xin Gao, Xiaolei Liu, Zidong Wang 等AAAI 2021 · 被引用 7 次
