ScaleFold: Reducing AlphaFold Initial Training Time to 10 Hours
Feiwen Zhu, Arkadiusz Nowaczynski, Rundong Li, Jie Xin, Yifei Song, Michal Marcinkiewicz, Sukru Burc Eryilmaz, Jun Yang, Michael Andersch
摘要
AlphaFold2 has been hailed as a breakthrough in protein folding. It can rapidly predict protein structures with lab-grade accuracy. However, its implementation does not include the necessary training code. OpenFold is the first trainable public reimplementation of AlphaFold. AlphaFold training procedure is prohibitively timeconsuming, and gets diminishing benefits from scaling to more compute resources. In this work, we conducted a comprehensive analysis on the AlphaFold training procedure based on Openfold, identified that inefficient communications and overhead-dominated computations were the key factors that prevented the AlphaFold training from effective scaling. We introduced ScaleFold, a systematic training method that incorporated optimizations specifically for these factors. ScaleFold successfully scaled the AlphaFold training to 2080 NVIDIA H100 GPUs with high resource utilization. In the MLPerf HPC v3.0 benchmark, ScaleFold finished the Open-Fold benchmark in 7.51 minutes, shown over 6× speedup than the baseline. For training the AlphaFold model from scratch, ScaleFold completed the pretraining in 10 hours, a significant improvement over the seven days required by the original AlphaFold pretraining baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- FastFold: Optimizing AlphaFold Training and Inference on GPU ClustersShenggan Cheng, Xuanlei Zhao, Guangyang Lu, Jiarui Fang 等PPoPP 2024 · 被引用 11 次
- SimpleFold: Folding Proteins is Simpler than You ThinkYuyang Wang, Jiarui Lu, Navdeep Jaitly, Joshua M. Susskind 等ICLR 2026 · 被引用 29 次
- AlphaFold Meets Flow Matching for Generating Protein EnsemblesBowen Jing, Bonnie Berger, Tommi S. JaakkolaICML 2024 · 被引用 229 次
- ExplainableFold: Understanding AlphaFold Prediction with Explainable AIJuntao Tan, Yongfeng ZhangKDD 2023 · 被引用 14 次
- Learning inverse folding from millions of predicted structuresChloe Hsu, Robert Verkuil, Jason Liu, Zeming Lin 等ICML 2022 · 被引用 560 次
