BlockGNN: Towards Efficient GNN Acceleration Using Block-Circulant Weight Matrices
Zhe Zhou, Bizhao Shi, Zhe Zhang, Yijin Guan, Guangyu Sun, Guojie Luo
Abstract
In recent years, Graph Neural Networks (GNNs) appear to be state-of-the-art algorithms for analyzing non-euclidean graph data. By applying deep-learning to extract high-level representations from graph structures, GNNs achieve extraordinary accuracy and great generalization ability in various tasks. However, with the ever-increasing graph sizes, more and more complicated GNN layers, and higher feature dimensions, the computational complexity of GNNs grows exponentially. How to inference GNNs in real time has become a challenging problem, especially for some resource-limited edge-computing platforms.
To tackle this challenge, we propose BlockGNN, a software-hardware co-design approach to realize efficient GNN acceleration. At the algorithm level, we propose to leverage block-circulant weight matrices to greatly reduce the complexity of various GNN models. At the hardware design level, we propose a pipelined CirCore architecture, which supports efficient block-circulant matrices computation. Basing on CirCore, we present a novel BlockGNN accelerator to compute various GNNs with low latency. Moreover, to determine the optimal configurations for diverse deployed tasks, we also introduce a performance and resource model that helps choose the optimal hardware parameters automatically. Comprehensive experiments on the ZC706 FPGA platform demonstrate that on various GNN tasks, BlockGNN achieves up to 8.3× speedup compared to the baseline HyGCN architecture and 111.9× energy reduction compared to the Intel Xeon CPU platform.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf2941e9-83e3-4ac2-99fd-c6a2e34b08f4Cited by top-tier papers4
- TGOpt: Redundancy-Aware Optimizations for Temporal Graph Attention NetworksYufeng Wang, Charith MendisPPoPP 2023 · 22 citations
- Lift: Exploiting Hybrid Stacked Memory for Energy-Efficient Processing of Graph Convolutional NetworksJiaxian Chen, Zhaoyu Zhong, Kaoyi Sun, Chenlin Ma et al.DAC 2023 · 10 citations
- RAHP: A Redundancy-aware Accelerator for High-performance Hypergraph Neural NetworkHui Yu, Yu Zhang, Ligang He, Yingqi Zhao et al.MICRO 2024 · 6 citations
- TaGNN: An Efficient Topology-aware Accelerator for High-performance Dynamic Graph Neural NetworkHui Yu, Yu Zhang, Ligang He, Bing Peng et al.SC 2025 · 2 citations
Builds on4
- HyGCN: A GCN Accelerator with Hybrid ArchitectureMingyu Yan, Lei Deng, Xing Hu, Ling Liang et al.HPCA 2020 · 338 citations
- GCN-RL Circuit Designer: Transferable Transistor Sizing with Graph Neural Networks and Reinforcement LearningHanrui Wang, Kuan Wang, Jiacheng Yang, Linxiao Shen et al.DAC 2020 · 326 citations
- AWB-GCN: A Graph Convolutional Network Accelerator with Runtime Workload RebalancingTong Geng, Ang Li, Runbin Shi, Chunshu Wu et al.MICRO 2020 · 299 citations
- Point-GNN: Graph Neural Network for 3D Object Detection in a Point CloudWeijing Shi, Raj RajkumarCVPR 2020
Related papers
- FlowGNN: A Dataflow Architecture for Real-Time Workload-Agnostic Graph Neural Network InferenceRishov Sarkar, Stefan Abi-Karam, Yuqi He, Lakshmi Sathidevi et al.HPCA 2023 · 100 citations
- Hardware-Aware Graph Neural Network Automated Design for Edge Computing PlatformsAo Zhou, Jianlei Yang, Yingjie Qi, Yumeng Shi et al.DAC 2023 · 15 citations
- CDA-GNN: A Chain-driven Accelerator for Efficient Asynchronous Graph Neural NetworkHui Yu, Yu Zhang, Ligang He, Donghao He et al.DAC 2024 · 4 citations
- Hardware Acceleration of Graph Neural NetworksAdam Auten, Matthew Tomei, Rakesh KumarDAC 2020 · 108 citations
- ReGNN: A Redundancy-Eliminated Graph Neural Networks AcceleratorCen Chen, Kenli Li, Yangfan Li, Xiaofeng ZouHPCA 2022 · 59 citations
