Tensor-Parallelism with Partially Synchronized Activations
Itay Lamprecht, Asaf Karnieli, Yair Hanani, Niv Giladi, Daniel Soudry
摘要
Training and inference of Large Language Models (LLMs) with tensor-parallelism requires substantial communication to synchronize activations. Our findings suggest that with a few minor adjustments to current practices, LLMs can be trained without fully synchronizing activations, reducing bandwidth demands. We name this "Communication-Aware Architecture for Tensor-parallelism" (CAAT-Net). We train a 7B parameter CAAT-Net model and show that tensor-parallel communication can be reduced by up to 50% with no significant drop in pretraining accuracy across nearly all evaluated benchmarks. We also experiment with smaller 130M and 1.1B models to show the robustness and scalability of our method. We find that, in some scenarios, validation loss can even improve when reducing communication. Finally, we demonstrate how CAAT-Net accelerates both training and inference workloads across various settings and model sizes. * This work was done while the author was at Intel. 39th Conference on Neural Information Processing Systems (NeurIPS 2025).
devices, the compute workload per device decreases, while communication payload per device remains relatively constant [7]. This means that the relative cost of communication grows as the compute-to-communication ratio decreases. In extreme cases, communication time can overcome computation time, and thus dominate the training process. For these reasons, even the largest language models are typically trained with a tensor-parallelism dimension of 8 [4], utilizing fast intra-node communication for the heavy all-reduce operations.
Improving tensor-parallelism efficiency is even more important given that the growth in compute power exceeds the growth in communication bandwidth [7]. This trend should further expose communication time in large-scale training. Minimizing tensor-parallelism communication can enable better hardware compute utilization and reduce overall training costs, especially when extending tensor-parallelism across nodes. This is in line with the recent trend of building multi-node systems with high bandwidth communication.
In traditional tensor-parallelism, the activation tensors after communication are identical on all devices, i.e., fully synchronized. In this work, we show that LLM training can converge without fully synchronizing the activation tensors in the all-reduce operation. This means that we allow activations to vary on different devices after communication. We show that without full synchronization, the current training practice needs to be slightly adjusted. Failing to do so leads to critical issues such as a mismatch between forward and backward passes and numerical issues, which often result in training divergence. Relying on this insight, we suggest the partial channel-reduce operation, in which only a subset of the channels in the hidden dimension of the activation tensors is reduced. Unlike regular all-reduce, activations are not identical on all devices after the partial channel-reduce operation. In the extreme case where no channels are synchronized in partial channel-reduce, the model resembles an ensemble, communicating only to compute the loss function and embeddings. In the case where all channels are reduced, the model is a vanilla transformer model. We introduce Communication-Aware Architecture for Tensor-parallelism (CAAT-Net) -a new model architecture that is tailored for tensor-parallelism by utilizing partial channel-reduce to decrease communication overhead. While CAAT-Net has a smaller communication overhead compared to an identical model with full all-reduce, the number of parameters and total compute stay the same.
We train a Llama2-7B model [8] with partial channel-reduce over 160B tokens and show that there is no significant degradation in nearly all evaluation benchmarks we tested, while reducing the communication payload by 50%. Furthermore, we train multiple variants of the 1.1B parameter TinyLlama model [9] and a smaller 130M parameter model. We study the effects of the number of synchronized channels and tensor-parallel dimension on accuracy. We find that a gradually reducing communication from full synchronization first yields a slight improvement in validation loss, but performance worsens when communication becomes too limited. Reducing the communication by 50% achieves either similar or slightly better validation loss for all models we tested. Finally, we show the training and inference speedup of our proposed method in various settings.
In summary, our contributions in this paper are as follows:
• We show that when using tensor-parallelism, LLMs can be trained without fully synchronizing activation tensors.
• We propose CAAT-Net, a novel architecture that significantly decreases communication traffic in training with tensor-parallelism by synchronizing only part of the activation tensors.
• We show that in various settings, CAAT-Net accelerates both training and inference, and achieves accuracy largely on par wit
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- TACO: Efficient Communication Compression of Intermediate Tensors for Scalable Tensor-Parallel LLM TrainingMan Liu, Xingchen Liu, Xingjian Tian, Bing Lu 等HPDC 2026 · 被引用 1 次
- TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode InferenceXiaojuan Tang, Fanxu Meng, Pingzhi Tang, Yuxuan Wang 等ASPLOS 2026
它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Breaking the computation and communication abstraction barrier in distributed machine learning workloadsAbhinav Jangda, Jun Huang, Guodong Liu, Amir Hossein Nodehi Sabet 等ASPLOS 2022 · 被引用 68 次
- At Stability's Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, Daniel SoudryICLR 2020 · 被引用 25 次
- Ladder-Residual: Parallelism-Aware Architecture for Accelerating Large Model Inference with Communication OverlappingMuru Zhang, Mayank Mishra, Zhongzhu Zhou, William Brandon 等ICML 2025
相关 Paper
- SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language ModelsHan-Byul Kim, Duc N. M. Hoang, Arnav Kundu, Mohammad Samragh 等ICML 2025
- PrimePar: Efficient Spatial-temporal Tensor Partitioning for Large Transformer Model TrainingHaoran Wang, Lei Wang, Haobo Xu, Ying Wang 等ASPLOS 2024 · 被引用 7 次
- Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU SystemsChen Zhang, Qijun Zhang, Zhuoshan Zhou, Yijia Diao 等HPCA 2026 · 被引用 1 次
- 3D-INA: An Exploration of Integrating In-Network Aggregation into 3D Parallelism for LLM TrainingHuifeng Xing, Hao Wang, Yinfan Hu, Xin Ai 等INFOCOM 2026
- WeiPipe: Weight Pipeline Parallelism for Communication-Effective Long-Context Large Model TrainingJunfeng Lin, Ziming Liu, Yang You, Jun Wang 等PPoPP 2025 · 被引用 5 次
