MLP-Offload: Multi-Level, Multi-Path Offloading for LLM Pre-training to Break the GPU Memory Wall
Avinash Maurya, M. Mustafa Rafique, Franck Cappello, Bogdan Nicolae
摘要
Training LLMs larger than the aggregated memory of multiple GPUs is increasingly necessary due to the faster growth of LLM sizes compared to GPU memory. To this end, multi-tier host memory or disk offloading techniques are proposed by state of art. Despite advanced asynchronous multi-tier read/write strategies, such offloading strategies result in significant I/O overheads in the critical path of training, resulting in slower iterations. To this end, we propose MLP-Offload, a novel multi-level, multi-path offloading engine specifically designed for optimizing LLM training on resource-constrained setups by mitigating I/O bottlenecks. We make several key observations that drive the design of MLP-Offload, such as I/O overheads during the update dominate the iteration time; I/O bandwidth of the third-level remote storage tier remains unutilized; and, contention due to concurrent offloading amplifies I/O bottlenecks. Driven by these insights, we design and implement MLP-Offload to offload the optimizer states across multiple tiers in a cache-efficient and concurrency-controlled fashion to mitigate I/O bottlenecks during the backward and update phases. Evaluations on models up to 280B parameters shows that MLP-Offload achieves 2.5 × faster iterations compared to the state-of-the-art LLM training runtimes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang 等ICLR 2023 · 被引用 295 次
- ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learningSamyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith 等SC 2021 · 被引用 254 次
- GEMINI: Fast Failure Recovery in Distributed Training with In-Memory CheckpointsZhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang 等SOSP 2023 · 被引用 61 次
相关 Paper
- DataStates-LLM: Lazy Asynchronous Checkpointing for Large Language ModelsAvinash Maurya, Robert Underwood, M. Mustafa Rafique, Franck Cappello 等HPDC 2024 · 被引用 34 次
- ACCO: Accumulate While You Communicate for Communication-Overlapped Sharded LLM TrainingAdel Nabli, Louis Fournier, Pierre Erbacher, Louis Serrano 等NeurIPS 2025 · 被引用 5 次
- PipeOffload: Improving Scalability of Pipeline Parallelism with Memory OptimizationXinyi Wan, Penghui Qi, Guangxing Huang, Min Lin 等ICML 2025
- Crimson: Collaborative Parameter Updates for Efficient Pipeline Training of Large Language ModelsYapeng Jiang, Wuhui Chen, Ganhong Huang, Yuzhou Huang 等EuroSys 2026
- Accelerating Multi-modal LLM Training with Adaptive Model Placement and ParallelizationYiming Yin, Shaohuai Shi, Qiang Wang, Xiaowen ChuINFOCOM 2026
