ThunderGNN: Unlocking Tensor Cores for Graph Neural Networks
YuAng Chen, Siyi Teng, Wenqi Weng, Jeffrey Xu Yu
Abstract
Graph Neural Networks (GNNs) have emerged as the state-of-the-art methodology for learning on graph-structured data, yet their performance is severely constrained by a fundamental mismatch between irregular graph sparsity and the rigid parallelism of modern hardware. While modern GPUs rely on Tensor Cores (TCs) to deliver massive computational throughput, these units demand strictly tiled, dense inputs—a requirement that conflicts with the extreme sparsity of real-world graphs. Existing frameworks fail to resolve this design conflict: they either fallback to legacy SIMT cores, leaving TCs underutilized, or they incur prohibitive memory bloat by forcing sparse data into dense tiles via excessive padding. To bridge this gap, we propose ThunderGNN, a hardware-aware acceleration system designed to reconcile graph irregularity with Tensor Core rigidity. ThunderGNN employs a unified co-design strategy comprising three key optimizations: (1) a sparsity-aware reordering algorithm that logically groups graph rows to maximize local density; (2) a Condensed Binarized Abstraction (CBA) storage layout that physically organizes the adjacency matrix into TC-aligned blocks without explicit zero-padding; and (3) a hardware-aware execution engine that efficiently streams compressed blocks directly into TCs. Extensive experiments on NVIDIA A100 GPUs demonstrate that ThunderGNN significantly outperforms state-of-the-art systems, achieving geometric mean speedups of 1.89× over DGL and 2.59× over PyG.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- GraphSAINT: Graph Sampling Based Inductive Learning MethodHanqing Zeng, Hongkuan Zhou, Ajitesh Srivastava, Rajgopal Kannan et al.ICLR 2020 · 1,155 citations
- Graph Condensation for Graph Neural NetworksWei Jin, Lingxiao Zhao, Shichang Zhang, Yozen Liu et al.ICLR 2022 · 203 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- Structure-free Graph Condensation: From Large-scale Graphs to Condensed Graph-free DataXin Zheng, Miao Zhang, Chunyang Chen, Quoc Viet Hung Nguyen et al.NeurIPS 2023 · 115 citations
- DistGNN: scalable distributed training for large-scale graph neural networksMd. Vasimuddin, Sanchit Misra, Guixiang Ma, Ramanarayan Mohanty et al.SC 2021 · 110 citations
Related papers
- Accelerating GNNs on GPU Sparse Tensor Cores through N: M Sparsity-Oriented Graph ReorderingJou-An Chen, Hsin-Hsuan Sung, Ruifeng Zhang, Ang Li et al.PPoPP 2025 · 6 citations
- TC-GNN: Bridging Sparse GNN Computation and Dense Tensor Cores on GPUsYuke Wang, Boyuan Feng, Zheng Wang, Guyue Huang et al.USENIX ATC 2023
- On Efficient Scaling of GNNs via IO-Aware Layers ImplementationsDaria Fomina, Daniil Krasylnikov, Alexey Boykov, Andrey Dolgovyazov et al.ICML 2026
- PruneGNN: Algorithm-Architecture Pruning Framework for Graph Neural Network AccelerationDeniz Gurevin, Mohsin Shan, Shaoyi Huang, Md Amit Hasan et al.HPCA 2024 · 28 citations
- MaxK-GNN: Extremely Fast GPU Kernel Design for Accelerating Graph Neural Networks TrainingHongwu Peng, Xi Xie, Kaustubh Shivdikar, Md Amit Hasan et al.ASPLOS 2024 · 32 citations
