SC2021Top-tier venue
Chimera: efficiently training large-scale neural networks with bidirectional pipelines
Shigang Li, Torsten Hoefler
Abstract
Training large deep learning models at scale is very challenging. This paper proposes Chimera, a novel pipeline parallelism scheme which combines bidirectional pipelines for efficiently training large-scale models. Chimera is a synchronous approach and therefore no loss of accuracy, which is more convergence-friendly than asynchronous approaches. Compared with the latest synchronous pipeline approach, Chimera reduces the number of bubbles by up to 50%; benefiting from the sophisticated scheduling of bidirectional pipelines, Chimera has a more balanced activation memory consumption. Evaluations are conducted on Transformer based language models. For a GPT-2 model with 1.3 billion parameters running on 2,048 GPU nodes of the Piz Daint supercomputer, Chimera improves the training throughput by 1.16x-2.34x over the state-of-the-art synchronous and asynchronous pipeline approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers48
- SmartMoE: Efficiently Training Sparsely-Activated Models through Combining Offline and Online ParallelizationMingshu Zhai, Jiaao He, Zixuan Ma, Zan Zong et al.USENIX ATC 2023 · 96 citations
- EE-LLM: Large-Scale Training and Inference of Early-Exit Large Language Models with 3D ParallelismYanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding et al.ICML 2024 · 73 citations
- Optimizing RLHF Training for Large Language Models with Stage FusionYinmin Zhong, Zili Zhang, Bingyang Wu, Shengyu Liu et al.NSDI 2025 · 64 citations
- Near-optimal sparse allreduce for distributed deep learningShigang Li, Torsten HoeflerPPoPP 2022 · 57 citations
- Extending the limit of molecular dynamics with ab initio accuracy to 10 billion atomsZhuoqiang Guo, Denghui Lu, Yujin Yan, Siyu Hu et al.PPoPP 2022 · 50 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu et al.ICML 2020 · 1,773 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- MixPipe: Efficient Bidirectional Pipeline Parallelism for Training Large-Scale ModelsWeigang Zhang, Biyu Zhou, Xuehai Tang, Zhaoxing Wang et al.DAC 2023 · 9 citations
- Hanayo: Harnessing Wave-like Pipeline Parallelism for Enhanced Large Model Training EfficiencyZiming Liu, Shenggan Cheng, Haotian Zhou, Yang YouSC 2023 · 35 citations
- Mario: Near Zero-cost Activation Checkpointing in Pipeline ParallelismWeijian Liu, Mingzhen Li, Guangming Tan, Weile JiaPPoPP 2025 · 4 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- BPipe: Memory-Balanced Pipeline Parallelism for Training Large Language ModelsTaebum Kim, Hyoungjoo Kim, Gyeong-In Yu, Byung-Gon ChunICML 2023 · 34 citations
