TLPGNN: A Lightweight Two-Level Parallelism Paradigm for Graph Neural Network Computation on GPU
Qiang Fu, Yuede Ji, H. Howie Huang
摘要
Graph Neural Networks (GNNs) are an emerging class of deep learning models on graphs, with many successful applications, such as, recommendation systems, drug discovery, and social network analysis. The GNN computation includes both regular neural network operations and general graph convolution operations, which take the majority of the total computation time. Though several recent works have been proposed to accelerate the computation for GNNs, they face the limitations of heavy pre-processing, low efficient atomic operations, and unnecessary kernel launches. In this paper, we design TLPGNN, a lightweight two-level parallelism paradigm for GNN computation. First, we conduct a systematic analysis on the hardware resource usage of GNN workloads to deeply understand the specialties of GNN workloads. With the insightful observations, we then divide the GNN computation into two levels, i.e., vertex parallelism for the first level and feature par- allelism for the second. Next, we employ a novel hybrid dynamic workload assignment to address the imbalanced workload distribution. Furthermore, we fuse the kernels to reduce the number of kernel launches and cache the frequently accessed data into registers to avoid unnecessary memory traffics. Together, TLPGNN is able to significantly outperform existing GNN computation systems, such as DGL, GNNAdivsor, and FeatGraph, by 5.6×, 7.7×, and 3.3×, respectively, on the average.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- PiPAD: Pipelined and Parallel Dynamic GNN Training on GPUsChunyang Wang, Desen Sun, Yuebin BaiPPoPP 2023 · 被引用 27 次
- Towards Lightweight Graph Neural Network Search with Curriculum Graph SparsificationBeini Xie, Heng Chang, Ziwei Zhang, Zeyang Zhang 等KDD 2024 · 被引用 5 次
- Hector: An Efficient Programming and Compilation Framework for Implementing Relational Graph Neural Networks in GPU ArchitecturesKun Wu, Mert Hidayetoglu, Xiang Song, Sitao Huang 等ASPLOS 2024 · 被引用 4 次
- Identifying and Analyzing Pitfalls in GNN SystemsYidong Gong, Arnab Kanti Tarafder, Saima Afrin, Pradeep KumarUSENIX ATC 2025 · 被引用 4 次
- Optimization of GNN Training Through Half-precisionArnab Kanti Tarafder, Yidong Gong, Pradeep KumarHPDC 2025
它引用的顶会 Paper8
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- MAGNN: Metapath Aggregated Graph Neural Network for Heterogeneous Graph EmbeddingXinyu Fu, Jiani Zhang, Ziqiao Meng, Irwin KingWWW 2020 · 被引用 1,149 次
- Are we really making much progress?: Revisiting, benchmarking and refining heterogeneous graph neural networksQingsong Lv, Ming Ding, Qiang Liu, Yuxiang Chen 等KDD 2021 · 被引用 249 次
- GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUsYuke Wang, Boyuan Feng, Gushu Li, Shuangchen Li 等OSDI 2021 · 被引用 163 次
- Understanding and bridging the gaps in current GNN performance optimizationsKezhao Huang, Jidong Zhai, Zhen Zheng, Youngmin Yi 等PPoPP 2021 · 被引用 87 次
相关 Paper
- FeatGraph: a flexible and efficient backend for graph neural network systemsYuwei Hu, Zihao Ye, Minjie Wang, Jiali Yu 等SC 2020 · 被引用 57 次
- NeutronTP: Load-Balanced Distributed Full-Graph GNN Training with Tensor ParallelismXin Ai, Hao Yuan, Zeyu Ling, Qiange Wang 等VLDB 2025 · 被引用 8 次
- MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU PlatformsYuke Wang, Boyuan Feng, Zheng Wang, Tong Geng 等OSDI 2023 · 被引用 46 次
- FlowGNN: A Dataflow Architecture for Real-Time Workload-Agnostic Graph Neural Network InferenceRishov Sarkar, Stefan Abi-Karam, Yuqi He, Lakshmi Sathidevi 等HPCA 2023 · 被引用 100 次
- WiseGraph: Optimizing GNN with Joint Workload Partition of Graph and OperationsKezhao Huang, Jidong Zhai, Liyan Zheng, Haojie Wang 等EuroSys 2024 · 被引用 11 次
