Chimera: Communication Fusion for Hybrid Parallelism in Large Language Models
Le Qin, Junwei Cui, Weilin Cai, Jiayi Huang
Abstract
Large Language Models (LLMs), exemplified by ChatGPT, have emerged as a predominant workload in current machine learning systems. To achieve efficient training and inference within the constraints of limited single-NPU memory capacity, deploying LLMs on multi-NPU systems typically adopt a hybrid approach that combines various parallelism patterns. This hybrid parallelism within LLMs introduces a significant amount of diverse collective communications. However, these frequent blocking communications impose a substantial burden on the multi-NPU systems. Overcoming the communication bottleneck is crucial to unlocking the potential of multi-NPU systems for efficient and scalable LLM processing.
This paper introduces Chimera, a communication fusion mechanism for hybrid parallelism in LLMs. We comprehensively analyze the communication processes of each LLM parallelism pattern, identify the communication redundancy in hybrid parallelism and eliminate redundancy by fusing adjacent communication operators during parallelism transformation. By reordering operations and generating redundancy-free communication operator, Chimera effectively mitigates communication bottleneck in hybrid LLM parallelism. Our results show that Chimera achieves 1.23-7.06× network bandwidth speedup. Additionally, the end-to-end performance of LLM forward pass and backward pass on different typical multi-NPU systems achieves respective 1.32-1.58× and 1.16-1.36× speedups on average compared with those without communication fusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 98451af5-df1b-4f55-baa8-f8d50d5db696Cited by top-tier papers2
- MoEntwine: Unleashing the Potential of Wafer-Scale Chips for Large-Scale Expert Parallel InferenceXinru Tang, Jingxiang Hou, Dingcheng Jiang, Taiquan Wei et al.HPCA 2026 · 4 citations
- Disdp: Disaggregating Compute, Network, and Storage for Model-Sharded Data-Parallel TrainingMo Sun, Zihan Yang, Changyue Liao, Yingtao Li et al.ISCA 2026 · 1 citation
Builds on28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
Related papers
- Chimera: An Analytical Optimizing Framework for Effective Compute-intensive Operators FusionSize Zheng, Siyuan Chen, Peidi Song, Renze Chen et al.HPCA 2023 · 46 citations
- Centauri: Enabling Efficient Scheduling for Communication-Computation Overlap in Large Model Training via Communication PartitioningChang Chen, Xiuhong Li, Qianchao Zhu, Jiangfei Duan et al.ASPLOS 2024 · 52 citations
- Chimera: efficiently training large-scale neural networks with bidirectional pipelinesShigang Li, Torsten HoeflerSC 2021 · 124 citations
- Synergistic Tensor and Pipeline ParallelismMengshi Qi, Jiaxuan Peng, Jie M. Zhang, Juan Zhu et al.NeurIPS 2025 · 2 citations
- ClusterFusion: Expanding Operator Fusion Scope for LLM Inference via Cluster-Level Collective PrimitiveXinhao Luo, Zihan Liu, Yangjie Zhou, Shihan Fang et al.NeurIPS 2025 · 9 citations
