Accelerating Multi-modal LLM Training with Adaptive Model Placement and Parallelization
Yiming Yin, Shaohuai Shi, Qiang Wang, Xiaowen Chu
摘要
Multi-modal large language models (MLLMs) have become a focal point of recent AI research, with numerous works investigating their architectures, training strategies, and real-world applications. Unlike uni-modal LLMs, MLLMs introduce heterogeneous modules with various workloads, including different input modalities, model sizes (parameters can be trainable or frozen), and architectures. This results in significantly reduced scaling efficiency when MLLMs are trained on GPU clusters. Current training systems either place all MLLM modules on the same GPUs to minimize communication overheads or separate different modules onto different GPUs for more flexible task scheduling. However, both approaches underestimate the impacts of model placement and parallelization, leading to sub-optimal training performance on GPU clusters. In this paper, we propose MoPPTrain to adaptively determine a model placement and parallelization strategy to minimize the MLLM training time. To achieve this, we first design a novel pipeline schedule for the colocated placement to enable communication tasks to be well overlapped with computation tasks. Then we theoretically analyze possible model placement and parallelization strategies for any given MLLM on a GPU cluster to formulate an optimization problem. Finally, we develop a cost model for the end-to-end iteration time, which allows us to derive the near-optimal solution to the optimization problem. We implement MoPPTrain atop Megatron-LM and conduct extensive experiments on a 40-GPU cluster. Experimental results show that our MoPPTrain outperforms state-of-the-art baselines (Megatron-LM, DistTrain, and Optimus) by 1.06× ∼ 4.2× in end-to-end training time.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- DistTrain: Addressing Model and Data Heterogeneity with Disaggregated Training for Multimodal Large Language ModelsZili Zhang, Yinmin Zhong, Yimin Jiang, Hanpeng Hu 等SIGCOMM 2025 · 被引用 15 次
- Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble ExploitationWeiqi Feng, Yangrui Chen, Shaoyu Wang, Yanghua Peng 等USENIX ATC 2025 · 被引用 34 次
- MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in ProductionChunyu Xue, Yangrui Chen, Jianyu Jiang, Ningxin Zheng 等EuroSys 2026
- FASOP: Fast yet Accurate Automated Search for Optimal Parallelization of Transformers on Heterogeneous GPU ClustersSunyeol Hwang, Eungyeong Lee, Hongseok Oh, Youngmin YiHPDC 2024 · 被引用 4 次
- DISTMM: Accelerating Distributed Multimodal Model TrainingJun Huang, Zhen Zhang, Shuai Zheng, Feng Qin 等NSDI 2024 · 被引用 37 次
