Extending the limit of molecular dynamics with ab initio accuracy to 10 billion atoms
Zhuoqiang Guo, Denghui Lu, Yujin Yan, Siyu Hu, Rongrong Liu, Guangming Tan, Ninghui Sun, Wanrun Jiang, Lijun Liu, Yixiao Chen, Linfeng Zhang, Mohan Chen
摘要
High-performance computing, together with a neural network model trained from data generated with first-principles methods, has greatly boosted applications of ab initio molecular dynamics in terms of spatial and temporal scales on modern supercomputers. Previous state-of-the-art can achieve 1 -- 2 nanoseconds molecular dynamics simulation per day for 100-million atoms on the entire Summit supercomputer. In this paper, we have significantly reduced the memory footprint and computational time by a comprehensive approach with both algorithmic and system innovations. The neural network model is compressed by model tabulation, kernel fusion, and redundancy removal. Then optimizations such as acceleration of customized kernel, tabulation of activation function, MPI+OpenMP parallelization are implemented on GPU and ARM architectures. Testing results of the copper system show that the optimized code can scale up to the entire machine of both Fugaku and Summit, and the corresponding system size can be extended by a factor of 134 to an unprecedented 17 billion atoms. The strong scaling of a 13.5-million atom copper system shows that the time-to-solution can be 7 times faster, reaching 11.2 nanoseconds per day. This work opens the door for unprecedentedly large-scale molecular dynamics simulations based on ab initio accuracy and can be potentially utilized in studying more realistic applications such as mechanical properties of metals, semiconductor devices, batteries, etc. The optimization techniques detailed in this paper also provide insight for relevant high-performance computing applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- DistMLIP: A Distributed Inference Platform for Machine Learning Interatomic PotentialsKevin Han, Bowen Deng, Amir Barati Farimani, Gerbrand CederICLR 2026 · 被引用 10 次
- Scaling Molecular Dynamics with ab initio Accuracy to 149 Nanoseconds per DayJianxiong Li, Boyang Li, Zhuoqiang Guo, Mingzhen Li 等SC 2024 · 被引用 9 次
- RLEKF: An Optimizer for Deep Potential with Ab Initio AccuracySiyu Hu, Wentao Zhang, Qiuchen Sha, Feng Pan 等AAAI 2023 · 被引用 5 次
- Deep Learning-Enabled Supercritical Flame Simulation at Detailed Chemistry and Real-Fluid Accuracy Towards Trillion-Cell ScaleZhuoqiang Guo, Runze Mao, Lijun Liu, Guangming Tan 等SC 2025 · 被引用 1 次
它引用的顶会 Paper7
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Chimera: efficiently training large-scale neural networks with bidirectional pipelinesShigang Li, Torsten HoeflerSC 2021 · 被引用 124 次
- A Mean Field Analysis Of Deep ResNet And Beyond: Towards Provably Optimization Via Overparameterization From DepthYiping Lu, Chao Ma, Yulong Lu, Jianfeng Lu 等ICML 2020 · 被引用 85 次
- Accelerating sparse DNN models without hardware-support via tile-wise sparsityCong Guo, Bo Yang Hsueh, Jingwen Leng, Yuxian Qiu 等SC 2020 · 被引用 65 次
- GEMS: GPU-enabled memory-aware model-parallelism system for distributed DNN trainingArpan Jain, Ammar Ahmad Awan, Asmaa M. Aljuhani, Jahanzeb Maqbool Hashmi 等SC 2020 · 被引用 48 次
相关 Paper
- Enabling large-scale correlated electronic structure calculations: scaling the RI-MP2 method on summitGiuseppe M. J. Barca, Jorge L. Galvez Vallejo, David L. Poole, Melisa Alkan 等SC 2021 · 被引用 18 次
- Training one DeePMD Model in Minutes: a Step towards Online LearningSiyu Hu, Tong Zhao, Qiuchen Sha, Enji Li 等PPoPP 2024 · 被引用 3 次
- Scaling Correlated Fragment Molecular Orbital Calculations on SummitGiuseppe M. J. Barca, Calum Snowdon, Jorge L. Galvez Vallejo, Fazeleh S. Kazemian 等SC 2022 · 被引用 25 次
- MISA-AKMC : Achieve Kinetic Monte Carlo Simulation of 20 Quadrillion Atoms on GPU ClustersShunde Li, Zhijie Pan, Ningming Nie, Jue Wang 等SC 2025 · 被引用 1 次
- TensorKMC: kinetic Monte Carlo simulation of 50 trillion atoms driven by deep learning on a new generation of Sunway supercomputerHonghui Shang, Xin Chen, Xingyu Gao, Rongfen Lin 等SC 2021 · 被引用 16 次
