JanusPipe: Efficient Pipeline Parallel Training for Machine Learning Interatomic Potentials
Hongyu Wang, Weijian Liu, Hongtao Xu, Yan Wang, Mingzhen Li, Weile Jia, Guangming Tan
摘要
Discovering atom-level phenomena requires molecular dynamics (MD) simulations with ab initio accuracy. Machine learning interatomic potentials (MLIPs) enable stable, high-accuracy MD simulations, and their models exhibit scaling-law trends similar to large language models. However, the lack of scalable and efficient distributed training systems for conservative MLIPs makes them difficult to scale. This is because conservative MLIPs inherently follow a double-backward execution pattern, which involves computing gradients during the forward pass. This pattern creates a mismatch with existing distributed training systems, especially for pipeline parallelism. Therefore, we present JanusPipe, an efficient 3D-parallel (PP/DP/GP) training system tailored for conservative MLIPs. It integrates SymFold to enable memory-efficient pipeline parallelism for conservative MLIPs, and WaveK to reduce pipeline bubbles by balancing the four-phase compute time. Experimental results on 32 GPUs show that JanusPipe improves throughput by and on average over 1F1B and Hanayo, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- PairNorm: Tackling Oversmoothing in GNNsLingxiao Zhao, Leman AkogluICLR 2020 · 被引用 590 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- EquiformerV2: Improved Equivariant Transformer for Scaling to Higher-Degree RepresentationsYi-Lun Liao, Brandon M. Wood, Abhishek Das, Tess E. SmidtICLR 2024 · 被引用 311 次
- UMA: A Family of Universal Models for AtomsBrandon M. Wood, Misko Dzamba, Xiang Fu, Meng Gao 等NeurIPS 2025 · 被引用 282 次
- Chimera: efficiently training large-scale neural networks with bidirectional pipelinesShigang Li, Torsten HoeflerSC 2021 · 被引用 124 次
相关 Paper
- DistMLIP: A Distributed Inference Platform for Machine Learning Interatomic PotentialsKevin Han, Bowen Deng, Amir Barati Farimani, Gerbrand CederICLR 2026 · 被引用 10 次
- Physics-Informed Weakly Supervised Learning For Interatomic PotentialsMakoto Takamoto, Viktor Zaverkin, Mathias NiepertICML 2025
- FlashTP: Fused, Sparsity-Aware Tensor Product for Machine Learning Interatomic PotentialsSeung Yul Lee, Hojoon Kim, Yutack Park, Dawoon Jeong 等ICML 2025
- Smooth Dynamic Cutoffs for Machine Learning Interatomic PotentialsKevin Han, Haolin Cong, Bowen Deng, Amir Barati FarimaniICML 2026 · 被引用 1 次
- A recipe for scalable attention-based ML potentials: unlocking long-range accuracy with all-to-all node attentionEric Qu, Brandon Wood, Aditi Krishnapriyan, Zachary UlissiICML 2026 · 被引用 14 次
