SuperMesh: Energy-Efficient Collective Communications for Accelerators
Sabuj Laskar, Pranati Majhi, Abdullah Muzahid, Eun Jung Kim
摘要
Chiplet-based Deep Neural Network (DNN) accelerators are a promising approach to meet the scalability demands of modern DNN models.Such accelerators usually utilize 2D mesh topologies.However, state-of-the-art collective communication algorithms often struggle within these topologies due to limited connectivity at border nodes, leading to communication bottlenecks and performance degradation.To address this challenge, we propose two novel topologies for chiplet-based accelerators aimed at improving collective communication performance and energy efficiency by integrating additional links parallel to the existing peripheral links of mesh topologies.The first proposed topology, SuperMesh Bi adds bidirectional links parallel to all peripheral links.In contrast, the second proposed topology, SuperMesh Alter , alternately adds bidirectional links parallel to the peripheral links, offering additional paths for data traversal.Both of the topologies adhere to a core principle-augmenting the outer region of mesh topologies with extra links to retain the original structure's latency and scalability, ensuring compatibility with chiplet-based accelerator designs and maintaining energy efficiency.To fully utilize these enhanced topologies, we co-designed pipelined collective algorithms for AllReduce, ReduceScatter, and AllGather.Our proposed algorithms and topologies achieve an average AllReduce speedup of 1.18-1.33×and a 1.77-2.22×speedup in ReduceScatter and AllGather compared to conventional 2D-mesh topologies.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- TidalMesh: Topology-Driven AllReduce Collective Communication for Mesh TopologyDongkyun Lim, John KimHPCA 2025 · 被引用 12 次
- SPACX: Silicon Photonics-based Scalable Chiplet Accelerator for DNN InferenceYuan Li, Ahmed Louri, Avinash KaranthHPCA 2022 · 被引用 32 次
- Scaling Deep-Learning Inference with Chiplet-based Architecture and Photonic InterconnectsYuan Li, Ahmed Louri, Avinash KaranthDAC 2021 · 被引用 21 次
- Enhancing Collective Communication in MCM Accelerators for Deep Learning TrainingSabuj Laskar, Pranati Majhi, Sungkeun Kim, Farabi Mahmud 等HPCA 2024 · 被引用 22 次
- FRED: A Wafer-scale Fabric for 3D Parallel DNN TrainingSaeed Rashidi, William Won, Sudarshan Srinivasan, Puneet Gupta 等ISCA 2025 · 被引用 8 次
