Ladder-Residual: Parallelism-Aware Architecture for Accelerating Large Model Inference with Communication Overlapping
Muru Zhang, Mayank Mishra, Zhongzhu Zhou, William Brandon, Jue Wang, Yoon Kim, Jonathan Ragan-Kelley, Shuaiwen Leon Song, Ben Athiwaratkun, Tri Dao
摘要
Large language model inference is both memoryintensive and time-consuming, often requiring distributed algorithms to efficiently scale. Various model parallelism strategies are used to partition computation across multiple devices, reducing memory load and computation time. However, using model parallelism requires communication of information between GPUs, which limits the gains obtained by scaling up the number of devices. We introduce Ladder Residual, a simple architectural modification applicable to all residualbased models that enables straightforward overlapping to hide the latency of communication. Our insight is that in addition to system optimizations, the model architecture can also be redesigned to decouple communication from computation. While Ladder Residual can allow communication-computation decoupling in conventional parallelism patterns, we focus on Tensor Parallelism in this paper, which is particularly bottlenecked by its heavy communication. For a Transformer model with 70B parameters, applying Ladder Residual to all its layers can achieve 29% end-to-end wall clock speedup at inference time with sharding over 8 devices. We train a 1.2B and 3.5B Ladder Residual based Transformer models from scratch and observe comparable performance to a standard dense transformer baseline. We also show that it is possible to convert parts of the Llama-3.1 8B model to our Ladder Residual architecture with minimal accuracy degradation by only retraining for 3B tokens.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Tensor-Parallelism with Partially Synchronized ActivationsItay Lamprecht, Asaf Karnieli, Yair Hanani, Niv Giladi 等NeurIPS 2025 · 被引用 6 次
- Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA ServingXinyu Wang, Jonas M. Kübler, Kailash Budhathoki, Yida Wang 等NeurIPS 2025
- TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode InferenceXiaojuan Tang, Fanxu Meng, Pingzhi Tang, Yuxuan Wang 等ASPLOS 2026
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 被引用 2,317 次
相关 Paper
- PrimePar: Efficient Spatial-temporal Tensor Partitioning for Large Transformer Model TrainingHaoran Wang, Lei Wang, Haobo Xu, Ying Wang 等ASPLOS 2024 · 被引用 7 次
- SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language ModelsHan-Byul Kim, Duc N. M. Hoang, Arnav Kundu, Mohammad Samragh 等ICML 2025
- TawPipe: Topology-Aware Weight Pipeline Parallelism for Accelerating Long-Context Large Models TrainingHouming Wu, Ling ChenAAAI 2026
- First Attentions Last: Better Exploiting First Attentions for Efficient Parallel TrainingGyudong Kim, Hyukju Na, Jin Kyu Kim, Hyunsung Jang 等NeurIPS 2025
- WeiPipe: Weight Pipeline Parallelism for Communication-Effective Long-Context Large Model TrainingJunfeng Lin, Ziming Liu, Yang You, Jun Wang 等PPoPP 2025 · 被引用 5 次
