PrimePar: Efficient Spatial-temporal Tensor Partitioning for Large Transformer Model Training
Haoran Wang, Lei Wang, Haobo Xu, Ying Wang, Yuming Li, Yinhe Han
摘要
With the rapid up-scaling of transformer-based large language models (LLM), training these models is becoming increasingly demanding on novel parallel training techniques. Tensor partitioning is an extensively researched parallel technique, encompassing data and model parallelism, and has a significant influence on LLM training performance. However, existing state-of-the-art parallel training systems are based on incomplete tensor partitioning space, where the distribution of partitioned sub-operators is limited to the spatial dimension. We discover that introducing the temporal dimension into tensor partitioning of LLM training instance provides extra opportunities to avoid collective communication across devices, saving memory space and also overlapping device-to-device communication with computation. In this paper, we propose a new tensor partition primitive that distributes sub-operators along both the spatial and temporal dimensions to further explore communication and memory overhead reduction over current solutions. This new primitive creates a broader parallelization space and leads to parallel solutions that achieve better training throughput with lower peak memory occupancy compared to state-of-the-art techniques. To efficiently deploy optimized parallel transformer model training to multiple devices, we further present an optimization algorithm that can find optimal parallel solutions from our spatial-temporal tensor partition space with acceptable search time. Our evaluation shows that our optimized tensor partitioning achieves up to 1.68 × training throughput with 69% peak memory occupancy compared to state-of-the-art distributed training systems when training LLMs. Upon scaling to 32 GPUs, the geo-mean speedup across benchmarks is 1.30 ×. When applied in 3D parallelism, up to 1.46 × training throughput can be achieved.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper4
- PipeThreader: Software-Defined Pipelining for Efficient DNN ExecutionYu Cheng, Lei Wang, Yining Shi, Yuqing Xia 等OSDI 2025 · 被引用 9 次
- Syncopate: Efficient Multi-GPU AI Kernels via Automatic Chunk-Centric Compute-Communication OverlapXinwei Qiang, Yue Guan, Zhengding Hu, Keren Zhou 等OSDI 2026 · 被引用 3 次
- MeshSlice: Efficient 2D Tensor Parallelism for Distributed DNN TrainingHyoungwook Nam, Gerasimos Gerogiannis, Josep TorrellasISCA 2025 · 被引用 3 次
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang 等HPCA 2026 · 被引用 2 次
相关 Paper
- Ladder-Residual: Parallelism-Aware Architecture for Accelerating Large Model Inference with Communication OverlappingMuru Zhang, Mayank Mishra, Zhongzhu Zhou, William Brandon 等ICML 2025
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Synergistic Tensor and Pipeline ParallelismMengshi Qi, Jiaxuan Peng, Jie M. Zhang, Juan Zhu 等NeurIPS 2025 · 被引用 2 次
- Tensor-Parallelism with Partially Synchronized ActivationsItay Lamprecht, Asaf Karnieli, Yair Hanani, Niv Giladi 等NeurIPS 2025 · 被引用 6 次
- FASOP: Fast yet Accurate Automated Search for Optimal Parallelization of Transformers on Heterogeneous GPU ClustersSunyeol Hwang, Eungyeong Lee, Hongseok Oh, Youngmin YiHPDC 2024 · 被引用 4 次
