MeshSlice: Efficient 2D Tensor Parallelism for Distributed DNN Training
Hyoungwook Nam, Gerasimos Gerogiannis, Josep Torrellas
Abstract
In distributed training of large DNN models, the scalability of onedimensional (1D) tensor parallelism (TP) is limited because of its high communication cost. 2D TP attains extra scalability and efficiency because it reduces communication relative to 1D TP. Unfortunately, existing algorithms for general matrix multiplication (GeMM) in 2D TP suffer from inefficiencies. Indeed, Cannon's algorithm incurs high traffic, SUMMA suffers from high synchronization overhead, and a 2D GeMM with collective communication operations does not overlap communication with computation. In addition, it is difficult to optimize the numerous parameters of 2D TP, including the dataflow, mesh shape, and sharding. As a result, human experts are needed to find efficient configurations of 2D TP. To address these problems, this paper proposes MeshSlice, a novel 2D GeMM algorithm for efficient 2D TP in distributed DNN training. The MeshSlice algorithm slices the collective communications into multiple partial collectives that allow overlapping communication with computation. As a result, MeshSlice hides most of the communication latency. We also present the MeshSlice LLM autotuner, which automates finding an efficient 2D GeMM dataflow configuration, mesh shape, and communication granularity for Large Language Model (LLM) training using analytical cost models. To evaluate MeshSlice, we simulate TPUv4 clusters training LLM models. We show that MeshSlice maintains good efficiency up to at least 256-way 2D TP. In a cluster of 256 TPUs, MeshSlice trains the GPT-3 and Megatron-NLG models 12.0% and 23.4% faster, respectively, than the state-of-the-art algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23eea22f-c043-418f-ab4e-c1f463490303Cited by top-tier papers1
Ask how each one uses itBuilds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning ModelsShibo Wang, Jinliang Wei, Amit Sabne, Andy Davis et al.ASPLOS 2023 · 64 citations
- Centauri: Enabling Efficient Scheduling for Communication-Computation Overlap in Large Model Training via Communication PartitioningChang Chen, Xiuhong Li, Qianchao Zhu, Jiangfei Duan et al.ASPLOS 2024 · 52 citations
- T3: Transparent Tracking & Triggering for Fine-grained Overlap of Compute & CollectivesSuchita Pati, Shaizeen Aga, Mahzabeen Islam, Nuwan Jayasena et al.ASPLOS 2024 · 13 citations
Related papers
- PrimePar: Efficient Spatial-temporal Tensor Partitioning for Large Transformer Model TrainingHaoran Wang, Lei Wang, Haobo Xu, Ying Wang et al.ASPLOS 2024 · 7 citations
- FASOP: Fast yet Accurate Automated Search for Optimal Parallelization of Transformers on Heterogeneous GPU ClustersSunyeol Hwang, Eungyeong Lee, Hongseok Oh, Youngmin YiHPDC 2024 · 4 citations
- Piper: Multidimensional Planner for DNN ParallelizationJakub Tarnawski, Deepak Narayanan, Amar PhanishayeeNeurIPS 2021 · 82 citations
- Accelerating Multi-modal LLM Training with Adaptive Model Placement and ParallelizationYiming Yin, Shaohuai Shi, Qiang Wang, Xiaowen ChuINFOCOM 2026
- TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language ModelsZhuohan Li, Siyuan Zhuang, Shiyuan Guo, Danyang Zhuo et al.ICML 2021 · 160 citations
