Accelerating neural network training: An analysis of the AlgoPerf competition
Priya Kasimbeg, Frank Schneider, Runa Eschenhagen, Juhan Bae, Chandramouli Shama Sastry, Mark Saroufim, Boyuan Feng, Less Wright, Edward Z. Yang, Zachary Nado, Sourabh Medapati, Philipp Hennig
Abstract
The goal of the ALGOPERF: TRAINING ALGORITHMS competition is to evaluate practical speed-ups in neural network training achieved solely by improving the underlying training algorithms. In the external tuning ruleset, submissions must provide workload-agnostic hyperparameter search spaces, while in the self-tuning ruleset they must be completely hyperparameter-free. In both rulesets, submissions are compared on time-to-result across multiple deep learning workloads, training on fixed hardware. This paper presents the inaugural ALGOPERF competition's results, which drew 18 diverse submissions from 10 teams. Our investigation reveals several key findings: (1) The winning submission in the external tuning ruleset, using DISTRIBUTED SHAMPOO, demonstrates the effectiveness of nondiagonal preconditioning over popular methods like ADAM, even when compared on wall-clock runtime. (2) The winning submission in the self-tuning ruleset, based on the SCHEDULE FREE ADAMW algorithm, demonstrates a new level of effectiveness for completely hyperparameter-free training algorithms. (3) The topscoring submissions were surprisingly robust to workload changes. We also discuss the engineering challenges encountered in ensuring a fair comparison between different training algorithms. These results highlight both the significant progress so far, and the considerable room for further improvements. Submissions. Submitted training algorithms must adhere to the fixed ALGOPERF API (Dahl et al., 2023, Sec. 4.2) and are limited to four submission functions: (1) update_params is responsible for 1 Note, the two speed-ups for SCHEDULE FREE ADAMW are computed across different sets of workloads. 2 github.com/mlcommons/algorithmic-efficiency/[...]/COMPETITION_RULES.md 3 github.com/mlcommons/algorithmic-efficiency/[...]/DOCUMENTATION.md
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers13
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 92 citations
- MuonBP: Faster Muon via Block-Periodic OrthogonalizationAhmed Khaled, Kaan Ozkara, Tao Yu, Mingyi Hong et al.ICLR 2026 · 35 citations
- Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its PreconditionerRuna Eschenhagen, Aaron Defazio, Tsung-Hsien Lee, Richard E. Turner et al.NeurIPS 2025 · 24 citations
- The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-NewtonNatalie Abreu, Nikhil Vyas, Sham M. Kakade, Depen MorwaniICLR 2026 · 21 citations
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei et al.NeurIPS 2025 · 17 citations
Builds on7
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- The Road Less ScheduledAaron Defazio, Xingyu Yang, Ahmed Khaled, Konstantin Mishchenko et al.NeurIPS 2024 · 208 citations
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 195 citations
- Prodigy: An Expeditiously Adaptive Parameter-Free LearnerKonstantin Mishchenko, Aaron DefazioICML 2024 · 131 citations
- Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language ModelsFrederik Kunstner, Alan Milligan, Robin Yadav, Mark Schmidt et al.NeurIPS 2024 · 100 citations
Related papers
- Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided PreconditioningHuan Li, Yiming Dong, Zhouchen LinICML 2026
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan et al.ICML 2026 · 1 citation
- When, Where and Why to Average Weights?Niccolò Ajroldi, Antonio Orvieto, Jonas GeipingICML 2025
- Dimension-Free Adaptive Subgradient Methods with Frequent DirectionsSifan Yang, Yuanyu Wan, Peijia Li, Yibo Wang et al.ICML 2025
- General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimizationKwangjun Ahn, Gagik Magakyan, Ashok CutkoskyICML 2025
