TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language Models
Zhuohan Li, Siyuan Zhuang, Shiyuan Guo, Danyang Zhuo, Hao Zhang, Dawn Song, Ion Stoica
摘要
Model parallelism has become a necessity for training modern large-scale deep language models. In this work, we identify a new and orthogonal dimension from existing model parallel approaches: it is possible to perform pipeline parallelism within a single training sequence for Transformer-based language models thanks to its autoregressive property. This enables a more fine-grained pipeline compared with previous work. With this key idea, we design TeraPipe, a high-performance token-level pipeline parallel algorithm for synchronous model-parallel training of Transformer-based language models. We develop a novel dynamic programming-based algorithm to calculate the optimal pipelining execution scheme given a specific model and cluster configuration. We show that TeraPipe can speed up the training by 5.0x for the largest GPT-3 model with 175 billion parameters on an AWS cluster with 48 p3.16xlarge instances compared with state-of-the-art model-parallel methods. The code for reproduction can be found at https://github.com/zhuohan123/terapipe
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- Mooncake: Trading More Storage for Less Computation - A KVCache-centric Architecture for Serving LLM ChatbotRuoyu Qin, Zheming Li, Weiran He, Jialei Cui 等FAST 2025 · 被引用 337 次
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu 等OSDI 2023 · 被引用 211 次
- Decentralized Training of Foundation Models in Heterogeneous EnvironmentsBinhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang 等NeurIPS 2022 · 被引用 157 次
- Skeleton-of-Thought: Prompting LLMs for Efficient Parallel GenerationXuefei Ning, Zinan Lin, Zixuan Zhou, Zifu Wang 等ICLR 2024 · 被引用 105 次
它引用的顶会 Paper5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 被引用 2,878 次
- Long Range Arena : A Benchmark for Efficient TransformersYi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen 等ICLR 2021 · 被引用 881 次
- ZeRO-Offload: Democratizing Billion-Scale Model TrainingJie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase 等USENIX ATC 2021 · 被引用 657 次
相关 Paper
- MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale ModelsWeigang Zhang, Biyu Zhou, Xuehai Tang, Zhaoxing Wang 等DAC 2023 · 被引用 9 次
- BPipe: Memory-Balanced Pipeline Parallelism for Training Large Language ModelsTaebum Kim, Hyoungjoo Kim, Gyeong-In Yu, Byung-Gon ChunICML 2023 · 被引用 34 次
- Chimera: efficiently training large-scale neural networks with bidirectional pipelinesShigang Li, Torsten HoeflerSC 2021 · 被引用 124 次
- FASOP: Fast yet Accurate Automated Search for Optimal Parallelization of Transformers on Heterogeneous GPU ClustersSunyeol Hwang, Eungyeong Lee, Hongseok Oh, Youngmin YiHPDC 2024 · 被引用 4 次
- Synergistic Tensor and Pipeline ParallelismMengshi Qi, Jiaxuan Peng, Jie M. Zhang, Juan Zhu 等NeurIPS 2025 · 被引用 2 次
