QGTC: accelerating quantized graph neural networks via GPU tensor core
Yuke Wang, Boyuan Feng, Yufei Ding
摘要
Over the most recent years, quantized graph neural network (QGNN) attracts lots of research and industry attention due to its high robustness and low computation and memory overhead. Unfortunately, the performance gains of QGNN have never been realized on modern GPU platforms. To this end, we propose the first Tensor Core (TC) based computing framework, QGTC, to support any-bitwidth computation for QGNNs on GPUs. We introduce a novel quantized low-bit arithmetic design based on the low-bit data representation and bit-decomposed computation. We craft a novel TC-tailored CUDA kernel design by incorporating 3D-stacked bit compression, zero-tile jumping, and non-zero tile reuse technique to improve the performance systematically. We incorporate an effective bandwidth-optimized subgraph packing strategy to maximize the transferring efficiency between CPU host and GPU device. We integrate QGTC with Pytorch for better programmability and extensibility. Extensive experiments demonstrate that QGTC can achieve evident inference speedup (on average 2.7X) compared with the state-of-the-art DGL framework across diverse settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- SparseTIR: Composable Abstractions for Sparse Compilation in Deep LearningZihao Ye, Ruihang Lai, Junru Shao, Tianqi Chen 等ASPLOS 2023 · 被引用 86 次
- DASP: Specific Dense Matrix Multiply-Accumulate Units Accelerated General Sparse Matrix-Vector MultiplicationYuechen Lu, Weifeng LiuSC 2023 · 被引用 37 次
- TGOpt: Redundancy-Aware Optimizations for Temporal Graph Attention NetworksYufeng Wang, Charith MendisPPoPP 2023 · 被引用 22 次
- DGC: Training Dynamic Graphs with Spatio-Temporal Non-Uniformity using Graph Partitioning by ChunksFahao Chen, Peng Li, Celimuge WuSIGMOD 2024 · 被引用 10 次
- TANGO: re-thinking quantization for graph neural network training on GPUsShiyang Chen, Da Zheng, Caiwen Ding, Chengying Huan 等SC 2023 · 被引用 10 次
它引用的顶会 Paper4
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Degree-Quant: Quantization-Aware Training for Graph Neural NetworksShyam Anil Tailor, Javier Fernández-Marqués, Nicholas Donald LaneICLR 2021 · 被引用 180 次
- A Broader Picture of Random-walk Based Graph EmbeddingZexi Huang, Arlei Silva, Ambuj K. SinghKDD 2021 · 被引用 42 次
- Binary Graph Neural NetworksMehdi Bahri, Gaétan Bahl, Stefanos ZafeiriouCVPR 2021
相关 Paper
- TC-GNN: Bridging Sparse GNN Computation and Dense Tensor Cores on GPUsYuke Wang, Boyuan Feng, Zheng Wang, Guyue Huang 等USENIX ATC 2023
- TenGraph: A Tensor-Based Graph Query EngineGuanghua Li, Hao Zhang, Xibo Sun, Qiong Luo 等VLDB 2024 · 被引用 4 次
- APNN-TC: accelerating arbitrary precision neural networks on ampere GPU tensor coresBoyuan Feng, Yuke Wang, Tong Geng, Ang Li 等SC 2021 · 被引用 36 次
- WholeGraph: A Fast Graph Neural Network Training Framework with Multi-GPU Distributed Shared Memory ArchitectureDongxu Yang, Junhong Liu, Jiaxing Qi, Junjie LaiSC 2022 · 被引用 12 次
- TGraph: A Tensor-centric Graph Processing FrameworkYongliang Zhang, Yuanyuan Zhu, Hao Zhang, Congli Gao 等SIGMOD 2025 · 被引用 1 次
