FastFold: Optimizing AlphaFold Training and Inference on GPU Clusters
Shenggan Cheng, Xuanlei Zhao, Guangyang Lu, Jiarui Fang, Tian Zheng, Ruidong Wu, Xiwen Zhang, Jian Peng, Yang You
摘要
Protein structure prediction helps to understand gene translation and protein function, which is of growing interest and importance in structural biology. The AlphaFold model, which used transformer architecture to achieve atomic-level accuracy in protein structure prediction, was a significant breakthrough. However, training and inference of AlphaFold model are challenging due to its high computation and memory cost. In this work, we present FastFold, an efficient implementation of AlphaFold for both training and inference. We propose Dynamic Axial Parallelism (DAP) as a novel model parallelism method. Additionally, we have implemented a series of low-level optimizations aimed at reducing communication, computation, and memory costs. These optimizations include Duality Async Operations, highly optimized kernels, and AutoChunk (an automated search algorithm finds the best chunk strategy to reduce memory peaks). Experimental results show that FastFold can efficiently scale to more GPUs using DAP and reduces overall training time from 11 days to 67 hours and achieves 7.5 9.5× speedup for long-sequence inference. Furthermore, AutoChunk can reduce memory cost by over 80% during inference by automatically partitioning the intermediate tensors during the computation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model TrainingZiming Liu, Shaoyu Wang, Shenggan Cheng, Zhongkai Zhao 等NeurIPS 2025 · 被引用 4 次
- LightNobel: Improving Sequence Length Limitation in Protein Structure Prediction Model via Adaptive Activation QuantizationSeunghee Han, Soongyu Choi, Joo-Young KimISCA 2025 · 被引用 1 次
相关 Paper
- ScaleFold: Reducing AlphaFold Initial Training Time to 10 HoursFeiwen Zhu, Arkadiusz Nowaczynski, Rundong Li, Jie Xin 等DAC 2024 · 被引用 6 次
- SimpleFold: Folding Proteins is Simpler than You ThinkYuyang Wang, Jiarui Lu, Navdeep Jaitly, Joshua M. Susskind 等ICLR 2026 · 被引用 29 次
- DCFold: Efficient Protein Structure Generation with Single Forward PassZhe Zhang, Yuanning Feng, Yuxuan Song, Keyue Qiu 等ICLR 2026 · 被引用 3 次
- Enabling Large Dynamic Neural Network Training with Learning-based Memory ManagementJie Ren, Dong Xu, Shuangyan Yang, Jiacheng Zhao 等HPCA 2024 · 被引用 11 次
- Sequence Parallelism: Long Sequence Training from System PerspectiveShenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li 等ACL 2023 · 被引用 29 次
