MPress: Democratizing Billion-Scale Model Training on Multi-GPU Servers via Memory-Saving Inter-Operator Parallelism
Quan Zhou, Haiquan Wang, Xiaoyan Yu, Cheng Li, Youhui Bai, Feng Yan, Yinlong Xu
Abstract
It remains challenging to train billion-scale DNN models on a single modern multi-GPU server due to the GPU memory wall. Unfortunately, existing memory-saving techniques such as GPU-CPU swap, recomputation, and ZeRO-Series come at the price of extra computation, communication overhead, or limited memory reduction.We present MPress, a new single-server multi-GPU system that breaks the GPU memory wall of billion-scale model training while minimizing extra cost. MPress first discusses the trade-offs of various memory-saving techniques and offers a holistic solution, which alternatively chooses the inter-operator parallelism with low cross-GPU communication traffics, and combines with recomputation and swap, to balance training performance and sustained model sizes. Additionally, MPress employs a novel, fast D2D swap technique, which simultaneously utilizes multiple high-bandwidth NVLink to swap tensors to light-load GPUs, based on a key observation that inter-operator parallel training may result in imbalanced GPU memory utilization and spare memory space from least used devices plus the high-end interconnects among them have the opportunity to support low-overhead swapping. Finally, we integrate MPress with PipeDream and DAPPLE, two representative inter-operator parallel training systems. Experimental results with two popular DNN models, Bert, and GPT, on two modern GPU servers from the DGX-1 and DGX-2 generation, equipped with 8 V100 or A100 cards, respectively, demonstrate that MPress significantly improves the training throughput over ZeRO-Series with the identical memory reduction, while being able to train larger models than the recomputation baseline.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers11
- Mobile Foundation Model as FirmwareJinliang Yuan, Chen Yang, Dongqi Cai, Shihe Wang et al.MobiCom 2024 · 40 citations
- AdaPipe: Optimizing Pipeline Parallelism with Adaptive Recomputation and PartitioningZhenbo Sun, Huanqi Cao, Yuanwei Wang, Guanyu Feng et al.ASPLOS 2024 · 28 citations
- CrossPipe: Towards Optimal Pipeline Schedules for Cross-Datacenter TrainingTiancheng Chen, Ales Kubicek, Langwen Huang, Torsten HoeflerUSENIX ATC 2025 · 20 citations
- MEPipe: Democratizing LLM Training with Memory-Efficient Slice-Level Pipeline Scheduling on Cost-Effective AcceleratorsZhenbo Sun, Shengqi Chen, Yuanwei Wang, Jian Sha et al.EuroSys 2025 · 7 citations
- FlexPipe: Maximizing Training Efficiency for Transformer-based Models with Variable-Length InputsHairui Zhao, Qi Tian, Hongliang Li, Zizhong ChenUSENIX ATC 2025 · 6 citations
Related papers
- Harmony: Overcoming the hurdles of GPU memory capacity to train massive DNN models on commodity serversYoujie Li, Amar Phanishayee, Derek Murray, Jakub Tarnawski et al.VLDB 2022 · 28 citations
- DAPPLE: a pipelined data parallel approach for training large modelsShiqing Fan, Yi Rong, Chen Meng, Zongyan Cao et al.PPoPP 2021 · 224 citations
- Group-based Interleaved Pipeline Parallelism for Large-scale DNN TrainingPengcheng Yang, Xiaoming Zhang, Wenpeng Zhang, Ming Yang et al.ICLR 2022 · 13 citations
- BPipe: Memory-Balanced Pipeline Parallelism for Training Large Language ModelsTaebum Kim, Hyoungjoo Kim, Gyeong-In Yu, Byung-Gon ChunICML 2023 · 34 citations
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith et al.SC 2021 · 254 citations
