Optimizing Data Distribution and Kernel Performance for Efficient Training of Chemistry Foundation Models: A Case Study with MACE
Jesun Sahariar Firoz, Franco Pellegrini, Mario Geiger, Darren Hsu, Jenna A. Bilbrey, Han-Yi Chou, Maximilian Stadler, Markus Höhnerbach, Tingyu Wang, Dejun Lin, Emine Küçükbenli, Henry W. Sprueill
摘要
Chemistry Foundation Models (CFMs) that leverage Graph Neural Networks (GNNs) operating on 3D molecular graph structures are becoming indispensable tools for computational chemists and materials scientists. These models facilitate the understanding of matter and the discovery of new molecules and materials. In contrast to GNNs operating on a large homogeneous graphs, GNNs used by CFMs process a large number of geometric graphs of varying sizes, requiring different optimization strategies than those developed for large homogeneous GNNs. This paper presents optimizations for two critical phases of CFM training: data distribution and model training, targeting MACE -a state-of-the-art CFM. We address the challenge of load balancing in data distribution by formulating it as a multi-objective bin packing problem. We propose an iterative algorithm that provides a highly effective, fast, and practical solution, ensuring efficient data distribution. For the training phase, we identify symmetric tensor contraction as the key computational kernel in MACE and optimize this kernel to improve the overall performance. Our combined approach of balanced data distribution and kernel optimization significantly enhances the training process of MACE. Experimental results demonstrate a substantial speedup, reducing per-epoch execution time for training from 12 to 2 minutes on 740 GPUs with a 2.6M sample dataset.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- MACE: Higher Order Equivariant Message Passing Neural Networks for Fast and Accurate Force FieldsIlyes Batatia, Dávid Péter Kovács, Gregor N. C. Simm, Christoph Ortner 等NeurIPS 2022 · 被引用 1,448 次
- E(n) Equivariant Graph Neural NetworksVictor Garcia Satorras, Emiel Hoogeboom, Max WellingICML 2021 · 被引用 1,432 次
- Directional Message Passing for Molecular GraphsJohannes Klicpera, Janek Groß, Stephan GünnemannICLR 2020 · 被引用 1,079 次
- GemNet: Universal Directional Graph Neural Networks for MoleculesJohannes Gasteiger, Florian Becker, Stephan GünnemannNeurIPS 2021 · 被引用 665 次
- Equiformer: Equivariant Graph Attention Transformer for 3D Atomistic GraphsYi-Lun Liao, Tess E. SmidtICLR 2023 · 被引用 65 次
相关 Paper
- Moment: Co-optimizing Physical Communication Topology and Data Placement for Multi-GPU Out-of-core GNN TrainingZuocheng Shi, Jie Sun, Ziyu Song, Mo Sun 等SC 2025 · 被引用 3 次
- Plexus: Taming Billion-edge Graphs with 3D Parallel Full-graph GNN TrainingAditya K. Ranjan, Siddharth Singh, Cunyang Wei, Abhinav BhateleSC 2025 · 被引用 1 次
- FAENet: Frame Averaging Equivariant GNN for Materials ModelingAlexandre Duval, Victor Schmidt, Alex Hernández-García, Santiago Miret 等ICML 2023 · 被引用 93 次
- On the Scalability of GNNs for Molecular GraphsMaciej Sypetkowski, Frederik Wenkel, Farimah Poursafaei, Nia Dickson 等NeurIPS 2024 · 被引用 58 次
- MIMOSA: Multi-constraint Molecule Sampling for Molecule OptimizationTianfan Fu, Cao Xiao, Xinhao Li, Lucas M. Glass 等AAAI 2021 · 被引用 94 次
