Overlap Communication with Dependent Computation via Decomposition in Large Deep Learning Models
Shibo Wang, Jinliang Wei, Amit Sabne, Andy Davis, Berkin Ilbeyi, Blake Hechtman, Dehao Chen, Karthik Srinivasa Murthy, Marcello Maggioni, Qiao Zhang, Sameer Kumar, Tongfei Guo
摘要
Large deep learning models have shown great potential with state-of-the-art results in many tasks. However, running these large models is quite challenging on an accelerator (GPU or TPU) because the on-device memory is too limited for the size of these models. Intra-layer model parallelism is an approach to address the issues by partitioning individual layers or operators across multiple devices in a distributed accelerator cluster. But, the data communications generated by intra-layer model parallelism can contribute to a significant proportion of the overall execution time and severely hurt the computational efficiency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper50
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan 等OSDI 2024 · 被引用 537 次
- Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble ExploitationWeiqi Feng, Yangrui Chen, Shaoyu Wang, Yanghua Peng 等USENIX ATC 2025 · 被引用 34 次
- Accelerating the Training of Large Language Models using Efficient Activation Rematerialization and Optimal Hybrid ParallelismTailing Yuan, Yuliang Liu, Xucheng Ye, Shenglong Zhang 等USENIX ATC 2024 · 被引用 34 次
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin 等NSDI 2025 · 被引用 27 次
相关 Paper
- Training Acceleration for Deep Neural Networks: A Hybrid Parallelization StrategyZihao Zeng, Chubo Liu, Zhuo Tang, Wanli Chang 等DAC 2021 · 被引用 13 次
- QuiltNet: efficient deep learning inference on multi-chip accelerators using model partitioningJongho Park, Hyukjun Kwon, Seowoo Kim, Junyoung Lee 等DAC 2022 · 被引用 6 次
- Alpa: Automating Inter- and Intra-Operator Parallelism for Distributed Deep LearningLianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang 等OSDI 2022 · 被引用 75 次
- Aceso: Efficient Parallel DNN Training through Iterative Bottleneck AlleviationGuodong Liu, Youshan Miao, Zhiqi Lin, Xiaoxiang Shi 等EuroSys 2024 · 被引用 16 次
- The ASPLOS 2025 / EuroSys 2025 Contest on Intra-Operator Parallelism for Distributed Deep LearningMichael D. Moffitt, Pratik FegadeASPLOS 2025
