Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
Anton Juerss, Vamsi Addanki, Stefan Schmid
摘要
AllReduce is a fundamental collective communication operation in distributed computing and a key performance bottleneck for large-scale training and inference. Its completion time is determined by the number of communication steps, which dominate latency-sensitive workloads, and the communication distance affecting both latency-and bandwidthbound regimes. Direct-connect topologies, such as Google's TPUv4 tori, are particularly prone to large communication distances due to limited bisection bandwidth.
In this paper, we present Trivance, a novel AllReduce algorithm that completes within log 3 𝑛 steps-a 50% improvement in comparison to Swing and Recursive Doubling, while reducing congestion compared to Bruck's algorithm by a factor of three and preserving bandwidth-optimality. Trivance exploits both transmission ports of a bidirectional ring within each step to triple the communication distance along both directions simultaneously. By performing joint reductions, Trivance improves both the number of steps and network congestion. We further show that Trivance extends naturally to multidimensional torus networks, retaining its latency advantage while achieving performance comparable to bandwidth-optimal algorithms for large AllReduce sizes.
Our packet-level SST simulation shows that Trivance improves state-of-the-art approaches by 5-30% for AllReduce sizes up to 8 MiB, in high-bandwidth settings up to 32 MiB and for 3D tori up to 128 MiB. Throughout the evaluation, Trivance remains the best-performing latency-optimal algorithm.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley 等SC 2021 · 被引用 576 次
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi 等NSDI 2023 · 被引用 215 次
- Synthesizing optimal collective algorithmsZixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi 等PPoPP 2021 · 被引用 64 次
- Swing: Short-cutting Rings for Higher Bandwidth AllreduceDaniele De Sensi, Tommaso Bonato, David Saam, Torsten HoeflerNSDI 2024 · 被引用 48 次
相关 Paper
- TidalMesh: Topology-Driven AllReduce Collective Communication for Mesh TopologyDongkyun Lim, John KimHPCA 2025 · 被引用 12 次
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid 等ISCA 2021 · 被引用 44 次
- Unified Communication Optimization Strategies for Sparse Triangular Solver on CPU and GPU ClustersYang Liu, Nan Ding, Piyush Sao, Samuel Williams 等SC 2023 · 被引用 8 次
- SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine LearningHans Kasan, Dennis Abts, Jungwook Choi, John KimMICRO 2025 · 被引用 1 次
- SuperMesh: Energy-Efficient Collective Communications for AcceleratorsSabuj Laskar, Pranati Majhi, Abdullah Muzahid, Eun Jung KimMICRO 2025 · 被引用 3 次
