Optimization of GNN Training Through Half-precision
Arnab Kanti Tarafder, Yidong Gong, Pradeep Kumar
Abstract
Recent trends in lower precision, e.g. half-precision floating point, training have shown improved system performance and reduced memory usage for Deep Learning while maintaining accuracy. However, current GNN systems cannot achieve such goals for GNN, as our analyses show that they massively underperform while showing abnormal accuracy when using half-precision. These systems suffer from under-utilization of hardware resources, poor training performance, and value overflow issues due to lowered precision. To mitigate this, we introduce HalfGNN, a half-precision based GNN system. HalfGNN proposes novel techniques: new vector operations for half-precision data types that improve data load and reduction performance, and discretized SpMM that overcomes the value overflow and natively provides workload balancing. Such techniques improve hardware utilization, reduce memory usage, and remove atomic writes. Evaluations show that HalfGNN achieves on average of 2.30× speedup in training time over DGL (float-based) for GAT, GCN, and GIN respectively while achieving similar accuracy, and saving 2.67× memory.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Identifying and Analyzing Pitfalls in GNN SystemsYidong Gong, Arnab Kanti Tarafder, Saima Afrin, Pradeep KumarUSENIX ATC 2025 · 4 citations
- ATLAS: Efficient Out-of-Core Inference for Billion-Scale Graph Neural NetworksPranjal Naman, Yogesh SimmhanHPDC 2026
Builds on13
- HyGCN: A GCN Accelerator with Hybrid ArchitectureMingyu Yan, Lei Deng, Xing Hu, Ling Liang et al.HPCA 2020 · 338 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
- GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUsYuke Wang, Boyuan Feng, Gushu Li, Shuangchen Li et al.OSDI 2021 · 163 citations
- GE-SpMM: general-purpose sparse matrix-matrix multiplication on GPUs for graph neural networksGuyue Huang, Guohao Dai, Yu Wang, Huazhong YangSC 2020 · 130 citations
- Relational Message Passing for Knowledge Graph CompletionHongwei Wang, Hongyu Ren, Jure LeskovecKDD 2021 · 109 citations
Related papers
- PruneGNN: Algorithm-Architecture Pruning Framework for Graph Neural Network AccelerationDeniz Gurevin, Mohsin Shan, Shaoyi Huang, Md Amit Hasan et al.HPCA 2024 · 28 citations
- TLPGNN: A Lightweight Two-Level Parallelism Paradigm for Graph Neural Network Computation on GPUQiang Fu, Yuede Ji, H. Howie HuangHPDC 2022 · 19 citations
- GNNOne: A Unified System Optimizations for GNN KernelsYidong Gong, Pradeep KumarHPDC 2024 · 3 citations
- ParGNN: A Scalable Graph Neural Network Training Framework on multi-GPUsJunyu Gu, Shunde Li, Rongqiang Cao, Jue Wang et al.DAC 2025
- XGNN: Boosting Multi-GPU GNN Training via Global GNN Memory StoreDahai Tang, Jiali Wang, Rong Chen, Lei Wang et al.VLDB 2024 · 13 citations
