SPD: Sync-Point Drop for Efficient Tensor Parallelism of Large Language Models
Han-Byul Kim, Duc N. M. Hoang, Arnav Kundu, Mohammad Samragh, Minsik Cho
Abstract
With the rapid expansion in the scale of large language models (LLMs), enabling efficient distributed inference across multiple computing units has become increasingly critical. However, communication overheads from popular distributed inference techniques such as Tensor Parallelism pose a significant challenge to achieve scalability and low latency. Therefore, we introduce a novel optimization technique, Sync-Point Drop (SPD), to reduce communication overheads in tensor parallelism by selectively dropping synchronization on attention outputs. In detail, we first propose a block design that allows execution to proceed without communication through SPD. Second, we apply different SPD strategies to attention blocks based on their sensitivity to the model accuracy. The proposed methods effectively alleviate communication bottlenecks while minimizing accuracy degradation during LLM inference, offering a scalable solution for diverse distributed environments: SPD offered about 20% overall inference latency reduction with < 1% accuracy regression for LLaMA2-70B inference over 8 GPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a1c0232-26f3-4090-99a4-0a631cdea314Cited by top-tier papers3
- Tensor-Parallelism with Partially Synchronized ActivationsItay Lamprecht, Asaf Karnieli, Yair Hanani, Niv Giladi et al.NeurIPS 2025 · 6 citations
- TPLA: Tensor Parallel Latent Attention for Efficient Disaggregated Prefill & Decode InferenceXiaojuan Tang, Fanxu Meng, Pingzhi Tang, Yuxuan Wang et al.ASPLOS 2026
- Kairox: Adaptive GPU-CPU Hybrid LLM Inference via Online Neuron BalancingYapeng Jiang, Minghao Gan, Zicong Hong, Wuhui Chen et al.OSDI 2026
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- QuIP: 2-Bit Quantization of Large Language Models With GuaranteesJerry Chee, Yaohui Cai, Volodymyr Kuleshov, Christopher De SaNeurIPS 2023 · 503 citations
Related papers
- Ladder-Residual: Parallelism-Aware Architecture for Accelerating Large Model Inference with Communication OverlappingMuru Zhang, Mayank Mishra, Zhongzhu Zhou, William Brandon et al.ICML 2025
- Shift Parallelism: Low-Latency, High-Throughput LLM Inference for Dynamic WorkloadsMert Hidayetoglu, Aurick Qiao, Michael Wyatt, Jeff Rasley et al.ASPLOS 2026 · 3 citations
- PrimePar: Efficient Spatial-temporal Tensor Partitioning for Large Transformer Model TrainingHaoran Wang, Lei Wang, Haobo Xu, Ying Wang et al.ASPLOS 2024 · 7 citations
- BurstEngine: An efficient distributed framework for training transformers On extremely Long sequences of over 1M tokensAo Sun, Weilin Zhao, Xu Han, Cheng Yang et al.SC 2025 · 1 citation
- First Attentions Last: Better Exploiting First Attentions for Efficient Parallel TrainingGyudong Kim, Hyukju Na, Jin Kyu Kim, Hyunsung Jang et al.NeurIPS 2025
