Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models
Shibo Wang, Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hechtman, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Qiao Zhang, Sameer Kumar, Tongfei Guo
Abstract
Large deep learning models have shown great potential with state-of-the-art results in many tasks. However, running these large models is quite challenging on an accelerator (GPU or TPU) because the on-device memory is too limited for the size of these models. Intra-layer model parallelism is an approach to address the issues by partitioning individual layers or operators across multiple devices in a distributed accelerator cluster. But, the data communications generated by intra-layer model parallelism can contribute to a significant proportion of the overall execution time and severely hurt the computational efficiency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bd4fa263-24c6-4bed-acb7-9f68ec24a580Cited by top-tier papers50
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski et al.ICML 2023 · 848 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
- Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble ExploitationWeiqi Feng, Yangrui Chen, Shaoyu Wang, Yanghua Peng et al.USENIX ATC 2025 · 34 citations
- Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid ParallelismTailing Yuan, Yuliang Liu, Xucheng Ye, Shenglong Zhang et al.USENIX ATC 2024 · 34 citations
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin et al.NSDI 2025 · 27 citations
Related papers
- Training Acceleration for Deep Neural Networks: A Hybrid Parallelization StrategyZihao Zeng, Chubo Liu, Zhuo Tang, Wanli Chang et al.DAC 2021 · 13 citations
- QuiltNet: efficient deep learning inference on multi-chip accelerators using model partitioningJongho Park, Hyukjun Kwon, Seowoo Kim, Junyoung Lee et al.DAC 2022 · 6 citations
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang et al.OSDI 2022 · 75 citations
- Aceso: Efficient Parallel DNN Training through Iterative Bottleneck AlleviationGuodong Liu, Youshan Miao, Zhiqi Lin, Xiaoxiang Shi et al.EuroSys 2024 · 16 citations
- The ASPLOS 2025 / EuroSys 2025 Contest on Intra-Operator Parallelism for Distributed Deep LearningMichael D. Moffitt, Pratik FegadeASPLOS 2025
