NetSparse: In-Network Acceleration of Distributed Sparse Kernels
Gerasimos Gerogiannis, Dimitrios Merkouriadis, Charles Block, Annus Zulfiqar, Filippos Tofalos, Muhammad Shahbaz, Josep Torrellas
摘要
Many hardware accelerators have been proposed to accelerate sparse computations. When these accelerators are placed in the nodes of a large cluster, distributed sparse applications become heavily communication-bound. Unfortunately, software solutions to optimize network communication are inefficient.
In this paper, we introduce novel hardware mechanisms to optimize network communication in distributed sparse computations. Our proposal, called NetSparse, consists of four mechanisms. Communication is offloaded to new processing units in the NIC that support efficient remote indexed gather operations, minimizing host-NIC communication. Moreover, these units have the ability to identify and eliminate redundant requests, which minimizes traffic. Further, new hardware modules in the NICs and switches concatenate multiple requests with the same destination node in a single packet, saving traffic and header overheads. Finally, switches are augmented with a hardware cache that stores fetched data from remote racks, making it available to all the nodes in the local rack on demand. Our evaluation on a simulated 128-node cluster with per-node sparse accelerators running sparse workloads reveals that NetSparse improves performance substantially. When the cluster uses traditional software-based communication, the workloads run only 3x faster than on a single-node system; when it is augmented with the NetSparse hardware, the workloads run 38x faster than on the single-node system-attaining more than half of the performance of an ideal system that has no communication overheads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor ComputationsNitish Kumar Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong 等HPCA 2020 · 被引用 121 次
- Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence DirectoriesJialin Li, Jacob Nelson, Ellis Michael, Xin Jin 等OSDI 2020 · 被引用 96 次
- Taurus: a data plane architecture for per-packet MLTushar Swamy, Alexander Rucker, Muhammad Shahbaz, Ishan Gaur 等ASPLOS 2022 · 被引用 94 次
- Flexagon: A Multi-dataflow Sparse-Sparse Matrix Multiplication Accelerator for Efficient DNN ProcessingFrancisco Muñoz-Martínez, Raveesh Garg, Michael Pellauer, José L. Abellán 等ASPLOS 2023 · 被引用 60 次
- Spada: Accelerating Sparse Matrix Multiplication with Adaptive DataflowZhiyao Li, Jiaxiang Li, Taijie Chen, Dimin Niu 等ASPLOS 2023 · 被引用 59 次
相关 Paper
- Nezha: An Efficient Distributed Graph Processing System on Heterogeneous HardwarePengjie Cui, Haotian Liu, Dong Jiang, Bo Tang 等SIGMOD 2025 · 被引用 2 次
- ABC-DIMM: Alleviating the Bottleneck of Communication in DIMM-based Near-Memory Processing with Inter-DIMM BroadcastWeiyi Sun, Zhaoshi Li, Shouyi Yin, Shaojun Wei 等ISCA 2021 · 被引用 42 次
- SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine LearningHans Kasan, Dennis Abts, Jungwook Choi, John KimMICRO 2025 · 被引用 1 次
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen 等NSDI 2021 · 被引用 359 次
- DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack CommunicationXu Zhang, Ke Liu, Yuan Hui, Xiaolong Zheng 等USENIX ATC 2025 · 被引用 5 次
