ADAHESSIAN: An Adaptive Second Order Optimizer for Machine Learning
Zhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa, Kurt Keutzer, Michael W. Mahoney
Abstract
We introduce AdaHessian, a second order stochastic optimization algorithm which dynamically incorporates the curvature of the loss function via ADAptive estimates of the Hessian. Second order algorithms are among the most powerful optimization algorithms with superior convergence properties as compared to first order methods such as SGD and Adam. The main disadvantage of traditional second order methods is their heavier per-iteration computation and poor accuracy as compared to first order methods. To address these, we incorporate several novel approaches in AdaHessian, including: (i) a fast Hutchinson based method to approximate the curvature matrix with low computational overhead; (ii) a root-mean-square exponential moving average to smooth out variations of the Hessian diagonal across different iterations; and (iii) a block diagonal averaging to reduce the variance of Hessian diagonal elements. We show that AdaHessian achieves new state-of-the-art results by a large margin as compared to other adaptive optimization methods, including variants of Adam. In particular, we perform extensive tests on CV, NLP, and recommendation system tasks and find that AdaHessian: (i) achieves 1.80%/1.45% higher accuracy on ResNets20/32 on Cifar10, and 5.55% higher accuracy on ImageNet as compared to Adam; (ii) outperforms AdamW for transformers by 0.13/0.33 BLEU score on IWSLT14/WMT14 and 2.7/1.0 PPL on PTB/Wikitext-103; (iii) outperforms AdamW for SqueezeBert by 0.41 points on GLUE; and (iv) achieves 0.032% better score than Adagrad for DLRM on the Criteo Ad Kaggle dataset. Importantly, we show that the cost per iteration of AdaHessian is comparable to first-order methods, and that it exhibits robustness towards its hyperparameters. The code for AdaHessian is open-sourced and publicly-available [1].
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 70645bc4-fffb-4de8-b508-28bc3ab954f1Cited by top-tier papers77
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda et al.NeurIPS 2020 · 697 citations
- Sophia: A Scalable Stochastic Second-order Optimizer for Language Model Pre-trainingHong Liu, Zhiyuan Li, David Leo Wright Hall, Percy Liang et al.ICLR 2024 · 264 citations
- The Right to be Forgotten in Federated Learning: An Efficient Realization with Rapid RetrainingYi Liu, Lei Xu, Xingliang Yuan, Cong Wang et al.INFOCOM 2022 · 189 citations
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 92 citations
- M-FAC: Efficient Matrix-Free Approximations of Second-Order InformationElias Frantar, Eldar Kurtic, Dan AlistarhNeurIPS 2021 · 69 citations
Builds on3
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- HAWQ: Hessian AWare Quantization of Neural Networks With Mixed-PrecisionZhen Dong, Zhewei Yao, Amir Gholami, Michael W. Mahoney et al.ICCV 2019 · 645 citations
- HAWQ-V2: Hessian Aware trace-Weighted Quantization of Neural NetworksZhen Dong, Zhewei Yao, Daiyaan Arfeen, Amir Gholami et al.NeurIPS 2020 · 434 citations
Related papers
- AdaFisher: Adaptive Second Order Optimization via Fisher InformationDamien Martins Gomes, Yanlei Zhang, Eugene Belilovsky, Guy Wolf et al.ICLR 2025
- Small Steps and Giant Leaps: Minimal Newton Solvers for Deep LearningJoão F. Henriques, Sébastien Ehrhardt, Samuel Albanie, Andrea VedaldiICCV 2019 · 23 citations
- Sassha: Sharpness-aware Adaptive Second-order Optimization with Stable Hessian ApproximationDahun Shin, Dongyeop Lee, Jinseok Chung, Namhoon LeeICML 2025
- A New Perspective on Shampoo's PreconditionerDepen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach et al.ICLR 2025
- ACMo: Angle-Calibrated Moment Methods for Stochastic OptimizationXunpeng Huang, Runxin Xu, Hao Zhou, Zhe Wang et al.AAAI 2021 · 2 citations
