FastFold: Optimizing AlphaFold Training and Inference on GPU Clusters
Shenggan Cheng, Xuanlei Zhao, Guangyang Lu, Jiarui Fang, Tian Zheng, Ruidong Wu, Xiwen Zhang, Jian Peng, Yang You
Abstract
Protein structure prediction helps to understand gene translation and protein function, which is of growing interest and importance in structural biology. The AlphaFold model, which used transformer architecture to achieve atomic-level accuracy in protein structure prediction, was a significant breakthrough. However, training and inference of AlphaFold model are challenging due to its high computation and memory cost. In this work, we present FastFold, an efficient implementation of AlphaFold for both training and inference. We propose Dynamic Axial Parallelism (DAP) as a novel model parallelism method. Additionally, we have implemented a series of low-level optimizations aimed at reducing communication, computation, and memory costs. These optimizations include Duality Async Operations, highly optimized kernels, and AutoChunk (an automated search algorithm finds the best chunk strategy to reduce memory peaks). Experimental results show that FastFold can efficiently scale to more GPUs using DAP and reduces overall training time from 11 days to 67 hours and achieves 7.5 9.5× speedup for long-sequence inference. Furthermore, AutoChunk can reduce memory cost by over 80% during inference by automatically partitioning the intermediate tensors during the computation.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get a2964581-f22d-4264-a12e-bafdadff725eCited by top-tier papers2
- StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model TrainingZiming Liu, Shaoyu Wang, Shenggan Cheng, Zhongkai Zhao et al.NeurIPS 2025 · 4 citations
- LightNobel: Improving Sequence Length Limitation in Protein Structure Prediction Model via Adaptive Activation QuantizationSeunghee Han, Soongyu Choi, Joo-Young KimISCA 2025 · 1 citation
Related papers
- ScaleFold: Reducing AlphaFold Initial Training Time to 10 HoursFeiwen Zhu, Arkadiusz Nowaczynski, Rundong Li, Jie Xin et al.DAC 2024 · 6 citations
- SimpleFold: Folding Proteins is Simpler than You ThinkYuyang Wang, Jiarui Lu, Navdeep Jaitly, Joshua M. Susskind et al.ICLR 2026 · 29 citations
- DCFold: Efficient Protein Structure Generation with Single Forward PassZhe Zhang, Yuanning Feng, Yuxuan Song, Keyue Qiu et al.ICLR 2026 · 3 citations
- Enabling Large Dynamic Neural Network Training with Learning-based Memory ManagementJie Ren, Dong Xu, Shuangyan Yang, Jiacheng Zhao et al.HPCA 2024 · 11 citations
- Sequence Parallelism: Long Sequence Training from System PerspectiveShenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li et al.ACL 2023 · 29 citations
