The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-Newton
Natalie Abreu, Nikhil Vyas, Sham M. Kakade, Depen Morwani
Abstract
Recent efforts to accelerate LLM pretraining have focused on computationally-efficient approximations that exploit second-order structure. This raises a key question for large-scale training: how much performance is forfeited by these approximations? To probe this question, we establish a practical upper bound on iteration complexity by applying full Gauss-Newton (GN) preconditioning to transformer models of up to 150M parameters. Our experiments show that full GN updates yield substantial gains over existing optimizers, achieving a 5.4x reduction in training iterations compared to strong baselines like SOAP and Muon. Furthermore, we find that a precise layerwise GN preconditioner, which ignores cross-layer information, nearly matches the performance of the full GN method. Collectively, our results suggest: (1) the GN approximation is highly effective for preconditioning, implying higher-order loss terms may not be critical for convergence speed; (2) the layerwise Hessian structure contains sufficient information to achieve most of these potential gains; and (3) a significant performance gap exists between current approximate methods and an idealized layerwise oracle. Code is available at: https://github.com/natalieabreu/full-gauss-newton .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9cb49642-6ef5-46fc-8cf2-ffb04563e30dCited by top-tier papers3
- RMNP: Row-Momentum Normalized Preconditioning for Scalable Matrix-Based OptimizationShenyang Deng, Zhuoli Ouyang, Tianyu Pang, Zihang Liu et al.ICML 2026 · 7 citations
- Mitigating Staleness in Asynchronous Pipeline Parallelism via Basis RotationHyunji Jung, Sungbin Shin, Namhoon LeeICML 2026 · 2 citations
- Geometric Convergence of Gauss–Newton for Neural Networks: Riemannian Geometry and Adaptive DampingSemih CayciICML 2026
Builds on6
- ADAHESSIAN: An Adaptive Second Order Optimizer for Machine LearningZhewei Yao, Amir Gholami, Sheng Shen, Mustafa Mustafa et al.AAAI 2021 · 358 citations
- Neural Tangents: Fast and Easy Infinite Neural Networks in PythonRoman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee et al.ICLR 2020 · 254 citations
- Gradient Alignment in Physics-informed Neural Networks: A Second-Order Optimization PerspectiveSifan Wang, Ananyae Kumar Bhartari, Bowen Li, Paris PerdikarisNeurIPS 2025 · 100 citations
- A New Perspective on Shampoo's PreconditionerDepen Morwani, Itai Shapira, Nikhil Vyas, Eran Malach et al.ICLR 2025
- Accelerating neural network training: An analysis of the AlgoPerf competitionPriya Kasimbeg, Frank Schneider, Runa Eschenhagen, Juhan Bae et al.ICLR 2025
Related papers
- AdamS: Momentum Itself Can Be A Normalizer for LLM Pretraining and Post-trainingHuishuai Zhang, Bohan Wang, Luoxin ChenEMNLP 2025 · 1 citation
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan et al.ICML 2026 · 1 citation
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei et al.NeurIPS 2025 · 17 citations
- SOAP: Improving and Stabilizing Shampoo using Adam for Language ModelingNikhil Vyas, Depen Morwani, Rosie Zhao, Itai Shapira et al.ICLR 2025
- Memory-Efficient LLM Pretraining via Minimalist Optimizer DesignAthanasios Glentis, Jiaxiang Li, Andi Han, Mingyi HongICML 2026 · 9 citations
