TidalMesh: Topology-Driven AllReduce Collective Communication for Mesh Topology
Dongkyun Lim, John Kim
Abstract
In deep learning workloads, collective communication across multiple nodes is a critical component in determining overall performance. AllReduce (as well as ReduceScatter and AllGather) is a commonly used collective communication for not only training but also inference. The performance of AllReduce depends on the algorithm utilized (or the “logical” topology) as well as the physical topology of the system that interconnects the nodes together. There has been many work on improving AllReduce performance but prior work have often been topologyaware approach where existing AllReduce algorithms were optimized for a given physical topology. In this work, we propose a topology-driven approach where the topology characteristics are exploited to propose a novel AllReduce collective communication algorithm; thus, the logical topology of the algorithm maps well to the physical topology. In particular, as 2D mesh topology is widely used in various scale-out systems, we propose TidalMesh AllReduce algorithm - a novel approach that exploits the inherent characteristics of the physical 2D mesh topology by pushing flows between the endpoint nodes, similar to a tidal wave, to achieve near-optimal performance for AllReduce. We propose how Sparse TidalMesh AllReduce minimizes bandwidth overhead of TidalMesh with no loss in performance. In addition, we demonstrate how collective communication unrolling can be exploited to enable “software pipelining” of collective communication while exploiting the unique opportunity of superimposing different phases of AllReduce. As a result, TidalMesh results in up to improvement in AllReduce performance across various deep learning models on a 64 -node 2D mesh, compared to the state-of-the-art while maintaining the simplicity of a logical ring algorithm.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers2
- TEMP: A Memory Efficient Physical-Aware Tensor Partition-Mapping Framework on Wafer-Scale ChipsHuizheng Wang, Taiquan Wei, Zichuan Wang, Dingcheng Jiang et al.HPCA 2026 · 2 citations
- FeLoG: Scalable and Efficient Distributed Graph Embedding with Feedback Loop MechanismPeng Fang, Arijit Khan, Ziqiang Wu, Zhenli Li et al.VLDB 2026
Related papers
- SuperMesh: Energy-Efficient Collective Communications for AcceleratorsSabuj Laskar, Pranati Majhi, Abdullah Muzahid, Eun Jung KimMICRO 2025 · 3 citations
- Logical/Physical Topology-Aware Collective Communication in Deep Learning TrainingJo Sanghoon, Hyojun Son, John KimHPCA 2023 · 23 citations
- Swing: Short-cutting Rings for Higher Bandwidth AllreduceDaniele De Sensi, Tommaso Bonato, David Saam, Torsten HoeflerNSDI 2024 · 48 citations
- Communication Algorithm-Architecture Co-Design for Distributed Deep LearningJiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid et al.ISCA 2021 · 44 citations
- Trivance: Latency-Optimal AllReduce by Shortcutting Multiport NetworksAnton Juerss, Vamsi Addanki, Stefan SchmidSIGCOMM 2026 · 1 citation
