Accelerating neural network training: An analysis of the AlgoPerf competition
Priya Kasimbeg, Frank Schneider, Runa Eschenhagen, Juhan Bae, Chandramouli Shama Sastry, Mark Saroufim, Boyuan Feng, Less Wright, Edward Z. Yang, Zachary Nado, Sourabh Medapati, Philipp Hennig
摘要
The goal of the ALGOPERF: TRAINING ALGORITHMS competition is to evaluate practical speed-ups in neural network training achieved solely by improving the underlying training algorithms. In the external tuning ruleset, submissions must provide workload-agnostic hyperparameter search spaces, while in the self-tuning ruleset they must be completely hyperparameter-free. In both rulesets, submissions are compared on time-to-result across multiple deep learning workloads, training on fixed hardware. This paper presents the inaugural ALGOPERF competition's results, which drew 18 diverse submissions from 10 teams. Our investigation reveals several key findings: (1) The winning submission in the external tuning ruleset, using DISTRIBUTED SHAMPOO, demonstrates the effectiveness of nondiagonal preconditioning over popular methods like ADAM, even when compared on wall-clock runtime. (2) The winning submission in the self-tuning ruleset, based on the SCHEDULE FREE ADAMW algorithm, demonstrates a new level of effectiveness for completely hyperparameter-free training algorithms. (3) The topscoring submissions were surprisingly robust to workload changes. We also discuss the engineering challenges encountered in ensuring a fair comparison between different training algorithms. These results highlight both the significant progress so far, and the considerable room for further improvements. Submissions. Submitted training algorithms must adhere to the fixed ALGOPERF API (Dahl et al., 2023, Sec. 4.2) and are limited to four submission functions: (1) update_params is responsible for 1 Note, the two speed-ups for SCHEDULE FREE ADAMW are computed across different sets of workloads. 2 github.com/mlcommons/algorithmic-efficiency/[...]/COMPETITION_RULES.md 3 github.com/mlcommons/algorithmic-efficiency/[...]/DOCUMENTATION.md
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Fantastic Pretraining Optimizers and Where to Find ThemKaiyue Wen, David Leo Wright Hall, Tengyu Ma, Percy LiangICLR 2026 · 被引用 92 次
- MuonBP: Faster Muon via Block-Periodic OrthogonalizationAhmed Khaled, Kaan Ozkara, Tao Yu, Mingyi Hong 等ICLR 2026 · 被引用 35 次
- Purifying Shampoo: Investigating Shampoo's Heuristics by Decomposing its PreconditionerRuna Eschenhagen, Aaron Defazio, Tsung-Hsien Lee, Richard E. Turner 等NeurIPS 2025 · 被引用 24 次
- The Potential of Second-Order Optimization for LLMs: A Study with Full Gauss-NewtonNatalie Abreu, Nikhil Vyas, Sham M. Kakade, Depen MorwaniICLR 2026 · 被引用 21 次
- Hyperparameter Transfer Enables Consistent Gains of Matrix-Preconditioned Optimizers Across ScalesShikai Qiu, Charlie Chen, Hoang Phan, Qi Lei 等NeurIPS 2025 · 被引用 17 次
它引用的顶会 Paper7
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein 等ASPLOS 2024 · 被引用 693 次
- The Road Less ScheduledAaron Defazio, Xingyu Yang, Ahmed Khaled, Konstantin Mishchenko 等NeurIPS 2024 · 被引用 208 次
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 被引用 195 次
- Prodigy: An Expeditiously Adaptive Parameter-Free LearnerKonstantin Mishchenko, Aaron DefazioICML 2024 · 被引用 131 次
- Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language ModelsFrederik Kunstner, Alan Milligan, Robin Yadav, Mark Schmidt 等NeurIPS 2024 · 被引用 100 次
相关 Paper
- Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-Sided and Two-Sided PreconditioningHuan Li, Yiming Dong, Zhouchen LinICML 2026
- DASH: Faster Shampoo via Batched Block Preconditioning and Efficient Inverse-Root SolversIonut-Vlad Modoranu, Philip Zmushko, Erik Schultheis, Mher Safaryan 等ICML 2026 · 被引用 1 次
- When, Where and Why to Average Weights?Niccolò Ajroldi, Antonio Orvieto, Jonas GeipingICML 2025
- Dimension-Free Adaptive Subgradient Methods with Frequent DirectionsSifan Yang, Yuanyu Wan, Peijia Li, Yibo Wang 等ICML 2025
- General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimizationKwangjun Ahn, Gagik Magakyan, Ashok CutkoskyICML 2025
