NetSparse: In-Network Acceleration of Distributed Sparse Kernels
Gerasimos Gerogiannis, Dimitrios Merkouriadis, Charles Block, Annus Zulfiqar, Filippos Tofalos, Muhammad Shahbaz, Josep Torrellas
Abstract
Many hardware accelerators have been proposed to accelerate sparse computations. When these accelerators are placed in the nodes of a large cluster, distributed sparse applications become heavily communication-bound. Unfortunately, software solutions to optimize network communication are inefficient.
In this paper, we introduce novel hardware mechanisms to optimize network communication in distributed sparse computations. Our proposal, called NetSparse, consists of four mechanisms. Communication is offloaded to new processing units in the NIC that support efficient remote indexed gather operations, minimizing host-NIC communication. Moreover, these units have the ability to identify and eliminate redundant requests, which minimizes traffic. Further, new hardware modules in the NICs and switches concatenate multiple requests with the same destination node in a single packet, saving traffic and header overheads. Finally, switches are augmented with a hardware cache that stores fetched data from remote racks, making it available to all the nodes in the local rack on demand. Our evaluation on a simulated 128-node cluster with per-node sparse accelerators running sparse workloads reveals that NetSparse improves performance substantially. When the cluster uses traditional software-based communication, the workloads run only 3x faster than on a single-node system; when it is augmented with the NetSparse hardware, the workloads run 38x faster than on the single-node system-attaining more than half of the performance of an ideal system that has no communication overheads.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4579d261-c7d9-45b4-8f22-ff6a22819553Builds on21
- Tensaurus: A Versatile Accelerator for Mixed Sparse-Dense Tensor ComputationsNitish Kumar Srivastava, Hanchen Jin, Shaden Smith, Hongbo Rong et al.HPCA 2020 · 121 citations
- Pegasus: Tolerating Skewed Workloads in Distributed Storage with In-Network Coherence DirectoriesJialin Li, Jacob Nelson, Ellis Michael, Xin Jin et al.OSDI 2020 · 96 citations
- Taurus: a data plane architecture for per-packet MLTushar Swamy, Alexander Rucker, Muhammad Shahbaz, Ishan Gaur et al.ASPLOS 2022 · 94 citations
- Flexagon: A Multi-dataflow Sparse-Sparse Matrix Multiplication Accelerator for Efficient DNN ProcessingFrancisco Muñoz-Martínez, Raveesh Garg, Michael Pellauer, José L. Abellán et al.ASPLOS 2023 · 60 citations
- Spada: Accelerating Sparse Matrix Multiplication with Adaptive DataflowZhiyao Li, Jiaxiang Li, Taijie Chen, Dimin Niu et al.ASPLOS 2023 · 59 citations
Related papers
- Nezha: An Efficient Distributed Graph Processing System on Heterogeneous HardwarePengjie Cui, Haotian Liu, Dong Jiang, Bo Tang et al.SIGMOD 2025 · 2 citations
- ABC-DIMM: Alleviating the Bottleneck of Communication in DIMM-based Near-Memory Processing with Inter-DIMM BroadcastWeiyi Sun, Zhaoshi Li, Shouyi Yin, Shaojun Wei et al.ISCA 2021 · 42 citations
- SkipReduce: (Interconnection) Network Sparsity to Accelerate Distributed Machine LearningHans Kasan, Dennis Abts, Jungwook Choi, John KimMICRO 2025 · 1 citation
- ATP: In-network Aggregation for Multi-tenant LearningChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen et al.NSDI 2021 · 359 citations
- DRack: A CXL-Disaggregated Rack Architecture to Boost Inter-Rack CommunicationXu Zhang, Ke Liu, Yuan Hui, Xiaolong Zheng et al.USENIX ATC 2025 · 5 citations
