Lune

MICRO2025顶会

SuperMesh: Energy-Efficient Collective Communications for Accelerators

Sabuj Laskar, Pranati Majhi, Abdullah Muzahid, Eun Jung Kim

2025年份
3被引次数
1顶会引用

摘要

Chiplet-based Deep Neural Network (DNN) accelerators are a promising approach to meet the scalability demands of modern DNN models.Such accelerators usually utilize 2D mesh topologies.However, state-of-the-art collective communication algorithms often struggle within these topologies due to limited connectivity at border nodes, leading to communication bottlenecks and performance degradation.To address this challenge, we propose two novel topologies for chiplet-based accelerators aimed at improving collective communication performance and energy efficiency by integrating additional links parallel to the existing peripheral links of mesh topologies.The first proposed topology, SuperMesh Bi adds bidirectional links parallel to all peripheral links.In contrast, the second proposed topology, SuperMesh Alter , alternately adds bidirectional links parallel to the peripheral links, offering additional paths for data traversal.Both of the topologies adhere to a core principle-augmenting the outer region of mesh topologies with extra links to retain the original structure's latency and scalability, ensuring compatibility with chiplet-based accelerator designs and maintaining energy efficiency.To fully utilize these enhanced topologies, we co-designed pipelined collective algorithms for AllReduce, ReduceScatter, and AllGather.Our proposed algorithms and topologies achieve an average AllReduce speedup of 1.18-1.33×and a 1.77-2.22×speedup in ReduceScatter and AllGather compared to conventional 2D-mesh topologies.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖