GNNOne: A Unified System Optimizations for GNN Kernels
Yidong Gong, Pradeep Kumar
Abstract
Graph Neural Networks (GNN) involve two basic sparse kernels, SDDMM and SpMM, on which all GNN models could be built. Prior works have explored piecemeal solutions by using different storage formats and computation paradigms, resulting in excess memory consumption, and have not yet realized their full potential. This paper, called GnnOne, studies these two basic sparse kernels in GPU and shows that they can be built on the same system design principle of data load being the limiting factor irrespective of their computing paradigms. Hence GnnOne presents a unified two-stage data-load design that provides greater performance through novel techniques of data-load balancing, data-load optimizations, and data-reuse. Such a unified design also enables the usage of a single sparse storage format to increase productivity, memory saving, and reduce maintenance. Evaluations show that the proposed system achieves an average speedup of 6.25× and 6.02× for SpMM and SDDMM over many prior works for different feature lengths. For GNN training, GnnOne achieves 2.01× average speedup over dgNN, 2.28× average speedup over DGL on 3 different GNN models.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- Identifying and Analyzing Pitfalls in GNN SystemsYidong Gong, Arnab Kanti Tarafder, Saima Afrin, Pradeep KumarUSENIX ATC 2025 · 4 citations
- Optimizing Data Distribution and Kernel Performance for Efficient Training of Chemistry Foundation Models: A Case Study with MACEJesun Sahariar Firoz, Franco Pellegrini, Mario Geiger, Darren Hsu et al.HPDC 2025 · 2 citations
- Optimization of GNN Training Through Half-precisionArnab Kanti Tarafder, Yidong Gong, Pradeep KumarHPDC 2025
Related papers
- XGNN: Boosting Multi-GPU GNN Training via Global GNN Memory StoreDahai Tang, Jiali Wang, Rong Chen, Lei Wang et al.VLDB 2024 · 13 citations
- MaxK-GNN: Extremely Fast GPU Kernel Design for Accelerating Graph Neural Networks TrainingHongwu Peng, Xi Xie, Kaustubh Shivdikar, Md Amit Hasan et al.ASPLOS 2024 · 32 citations
- GE-SpMM: general-purpose sparse matrix-matrix multiplication on GPUs for graph neural networksGuyue Huang, Guohao Dai, Yu Wang, Huazhong YangSC 2020 · 130 citations
- MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU PlatformsYuke Wang, Boyuan Feng, Zheng Wang, Tong Geng et al.OSDI 2023 · 46 citations
- TLPGNN: A Lightweight Two-Level Parallelism Paradigm for Graph Neural Network Computation on GPUQiang Fu, Yuede Ji, H. Howie HuangHPDC 2022 · 19 citations
