SuperMesh: Energy-Efficient Collective Communications for Accelerators
Sabuj Laskar, Pranati Majhi, Abdullah Muzahid, Eun Jung Kim
Abstract
Chiplet-based Deep Neural Network (DNN) accelerators are a promising approach to meet the scalability demands of modern DNN models.Such accelerators usually utilize 2D mesh topologies.However, state-of-the-art collective communication algorithms often struggle within these topologies due to limited connectivity at border nodes, leading to communication bottlenecks and performance degradation.To address this challenge, we propose two novel topologies for chiplet-based accelerators aimed at improving collective communication performance and energy efficiency by integrating additional links parallel to the existing peripheral links of mesh topologies.The first proposed topology, SuperMesh Bi adds bidirectional links parallel to all peripheral links.In contrast, the second proposed topology, SuperMesh Alter , alternately adds bidirectional links parallel to the peripheral links, offering additional paths for data traversal.Both of the topologies adhere to a core principle-augmenting the outer region of mesh topologies with extra links to retain the original structure's latency and scalability, ensuring compatibility with chiplet-based accelerator designs and maintaining energy efficiency.To fully utilize these enhanced topologies, we co-designed pipelined collective algorithms for AllReduce, ReduceScatter, and AllGather.Our proposed algorithms and topologies achieve an average AllReduce speedup of 1.18-1.33×and a 1.77-2.22×speedup in ReduceScatter and AllGather compared to conventional 2D-mesh topologies.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 568e96a8-3434-4927-a9b2-8ba96326c0f8Cited by top-tier papers1
Ask how each one uses itRelated papers
- TidalMesh: Topology-Driven AllReduce Collective Communication for Mesh TopologyDongkyun Lim, John KimHPCA 2025 · 12 citations
- SPACX: Silicon Photonics-based Scalable Chiplet Accelerator for DNN InferenceYuan Li, Ahmed Louri, Avinash KaranthHPCA 2022 · 32 citations
- Scaling Deep-Learning Inference with Chiplet-based Architecture and Photonic InterconnectsYuan Li, Ahmed Louri, Avinash KaranthDAC 2021 · 21 citations
- Enhancing Collective Communication in MCM Accelerators for Deep Learning TrainingSabuj Laskar, Pranati Majhi, Sungkeun Kim, Farabi Mahmud et al.HPCA 2024 · 22 citations
- FRED: A Wafer-scale Fabric for 3D Parallel DNN TrainingSaeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta et al.ISCA 2025 · 8 citations
