GEMS: GPU-enabled memory-aware model-parallelism system for distributed DNN training
Arpan Jain, Ammar Ahmad Awan, Asmaa M. Aljuhani, Jahanzeb Maqbool Hashmi, Quentin G. Anthony, Hari Subramoni, Dhabaleswar K. Panda, Raghu Machiraju, Anil Parwani
摘要
Data-parallelism has become an established paradigm to train DNNs that fit inside GPU memory on large-scale HPC systems. However, model-parallelism is required to train out-of-core DNNs. In this paper, we deal with emerging requirements brought forward by very large DNNs being trained using high-resolution images common in digital pathology. To address these, we propose, design, and implement GEMS; a GPU-Enabled Memory-Aware Model-Parallelism System. We present several design schemes like GEMS-MAST, GEMS-MASTER, and GEMS-Hybrid that offer excellent speedups over state-of-the-art systems like Mesh-TensorFlow and FlexFlow. Furthermore, we combine model-parallelism and data-parallelism to train a 1000-1ayer ResNet-lk model using 1,024 Volta V100 GPUs with 97.32% scaling-efficiency. For the real-world histopathology whole-slide-image (WSI) of 100,000 x 100,000 pixels, we train custom ResNet-110-v2 on image tiles of size 1024 x 1024 and reduce the training time from seven hours to 28 minutes.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- Chimera: efficiently training large-scale neural networks with bidirectional pipelinesShigang Li, Torsten HoeflerSC 2021 · 被引用 124 次
- Extending the limit of molecular dynamics with ab initio accuracy to 10 billion atomsZhuoqiang Guo, Denghui Lu, Yujin Yan, Siyu Hu 等PPoPP 2022 · 被引用 50 次
- Pipeline Parallelism with Controllable MemoryPenghui Qi, Xinyi Wan, Nyamdavaa Amar, Min LinNeurIPS 2024 · 被引用 25 次
- Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelismSaar Eliad, Ido Hakimi, Alon De Jagger, Mark Silberstein 等USENIX ATC 2021 · 被引用 24 次
- GraphPipe: Improving Performance and Scalability of DNN Training with Graph Pipeline ParallelismByungsoo Jeon, Mengdi Wu, Shiyi Cao, Sunghyun Kim 等ASPLOS 2025 · 被引用 10 次
相关 Paper
- A Scalable Distributed Framework for Multimodal GigaVoxel Image RegistrationRohit Jena, Vedant Zope, Pratik Chaudhari, James GeeICLR 2026
- FlashMoE: Fast Distributed MoE in a Single KernelOsayamen Jonathan Aimuyo, Byungsoo Oh, Rachee SinghNeurIPS 2025 · 被引用 22 次
- DAPPLE: a pipelined data parallel approach for training large modelsShiqing Fan, Yi Rong, Chen Meng, Zongyan Cao 等PPoPP 2021 · 被引用 224 次
- MPress: Democratizing Billion-Scale Model Training on Multi-GPU Servers via Memory-Saving Inter-Operator ParallelismQuan Zhou, Haiquan Wang, Xiaoyan Yu, Cheng Li 等HPCA 2023 · 被引用 20 次
- Tensor Movement Orchestration in Multi-GPU Training SystemsShao-Fu Lin, Yi-Jung Chen, Hsiang-Yun Cheng, Chia-Lin YangHPCA 2023 · 被引用 5 次
