Swing: Short-cutting Rings for Higher Bandwidth Allreduce
Daniele De Sensi, Tommaso Bonato, David Saam, Torsten Hoefler
摘要
The allreduce collective operation accounts for a significant fraction of the runtime of workloads running on distributed systems. One factor determining its performance is the number of hops between communicating nodes, especially on networks like torus, where a higher number of hops implies multiple messages being forwarded on the same link, thus reducing the allreduce bandwidth. Torus networks are widely used on systems optimized for machine learning workloads (e.g., Google TPUs and Amazon Trainium devices), as well as on some of the Top500 supercomputers. To improve allreduce performance on torus networks we introduce Swing, a new algorithm that reduces the number of hops between communicating nodes by swinging between torus directions. Our analysis and experimental evaluation show that Swing outperforms by up to 3x existing allreduce algorithms for vectors ranging from 32B to 128MiB, on different types of torus and torus-like topologies, regardless of their shape and size.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- AutoCCL: Automated Collective Communication Tuning for Accelerating Distributed and Parallel DNN TrainingGuanbin Xu, Zhihao Le, Yinhe Chen, Zhiqi Lin 等NSDI 2025 · 被引用 27 次
- Holmes: Localizing Irregularities in LLM Training with Mega-scale GPU ClustersZhiyi Yao, Pengbo Hu, Congcong Miao, Xuya Jia 等NSDI 2025 · 被引用 23 次
- MCCS: A Service-based Approach to Collective Communication for Multi-Tenant CloudYongji Wu, Yechen Xu, Jingrong Chen, Zhaodong Wang 等SIGCOMM 2024 · 被引用 15 次
- ResCCL: Resource-Efficient Scheduling for Collective CommunicationTongrui Liu, Chenyang Hei, Fuliang Li, Chengxi Gao 等SIGCOMM 2025 · 被引用 11 次
- Switch-Less Dragonfly on Wafers: A Scalable Interconnection Architecture based on Wafer-Scale IntegrationYinxiao Feng, Kaisheng MaSC 2024 · 被引用 10 次
它引用的顶会 Paper10
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi 等NSDI 2023 · 被引用 215 次
- An in-depth analysis of the slingshot interconnectDaniele De Sensi, Salvatore Di Girolamo, Kim H. McMahon, Duncan Roweth 等SC 2020 · 被引用 122 次
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini 等SIGCOMM 2021 · 被引用 120 次
- An In-Network Architecture for Accelerating Shared-Memory Multiprocessor CollectivesBenjamin Klenk, Nan Jiang, Greg Thorson, Larry DennisonISCA 2020 · 被引用 67 次
相关 Paper
- Trivance: Latency-Optimal AllReduce by Shortcutting Multiport NetworksAnton Juerss, Vamsi Addanki, Stefan SchmidSIGCOMM 2026 · 被引用 1 次
- TidalMesh: Topology-Driven AllReduce Collective Communication for Mesh TopologyDongkyun Lim, John KimHPCA 2025 · 被引用 12 次
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid 等ISCA 2021 · 被引用 44 次
- Efficient all-to-all Collective Communication Schedules for Direct-connect TopologiesPrithwish Basu, Liangyu Zhao, Jason Fantl, Siddharth Pal 等HPDC 2024 · 被引用 7 次
- FlexReduce: Flexible All-reduce for Distributed Deep Learning on Asymmetric Network TopologyJinho Lee, Inseok Hwang, Soham Shah, Minsik ChoDAC 2020 · 被引用 21 次
