TaGNN: An Efficient Topology-aware Accelerator for High-performance Dynamic Graph Neural Network
Hui Yu, Yu Zhang, Ligang He, Bing Peng, Jin Zhao, Zixiao Wang, Hao Qi, Hai Jin
摘要
Dynamic Graph Neural Networks (DGNNs) have become powerful tools for analyzing continuously evolving graph data, combining Graph Neural Network (GNN) models to extract structural information and Recurrent Neural Network (RNN) models to capture temporal semantics across snapshots. However, despite extensive research, existing DGNN solutions still face significant limitations, particularly low data parallelism caused by their snapshot-by-snapshot execution. This sequential paradigm exacerbates memory contention due to irregular, repeated vertex feature accesses and enforces strict temporal dependencies. In this paper, we propose TaGNN, an efficient topology-aware DGNN accelerator that addresses these performance bottlenecks. Specifically, we present a topology-aware concurrent execution approach into the accelerator design that calculates the final features of affected vertices while ensuring that unaffected vertices are loaded and computed only once per layer across multiple snapshots, maximizing data parallelism while minimizing memory usage. TaGNN employs a cache-friendly storage format that compactly organizes affected vertices across multiple snapshots by their timestamps and topological characteristics, reducing indexing overhead and enhancing data locality. In addition, TaGNN further proposes a similarity-aware cell skipping strategy to alleviate the stringent temporal data dependencies. It selectively reuses the RNN results from the previous snapshot to bypass RNN operations in the current snapshot when the output features of the GNN module across two consecutive snapshots are similar, achieving significant efficiency gains with minimal accuracy loss. We have implemented and assessed TaGNN on a Xilinx Alveo U280 FPGA card. Experimental results show that TaGNN achieves average speedups of 535.2x and 84.3x, and energy savings of 742.6x and 104.9x over state-of-the-art software DGNNs on Intel Xeon CPUs and NVIDIA A100 GPUs, respectively. Compared to leading DGNN accelerators (i.e., DGNN-Booster, E-DGCN, and Cambricon-DG), TaGNN delivers average speedups of 13.5x, 10.2x, and 6.5x, and energy savings of 15.9x, 11.7x, and 7.8x, respectively.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- EvolveGCN: Evolving Graph Convolutional Networks for Dynamic GraphsAldo Pareja, Giacomo Domeniconi, Jie Chen, Tengfei Ma 等AAAI 2020 · 被引用 1,429 次
- HyGCN: A GCN Accelerator with Hybrid ArchitectureMingyu Yan, Lei Deng, Xing Hu, Ling Liang 等HPCA 2020 · 被引用 338 次
- ReGNN: A Redundancy-Eliminated Graph Neural Networks AcceleratorCen Chen, Kenli Li, Yangfan Li, Xiaofeng ZouHPCA 2022 · 被引用 59 次
- P-OPT: Practical Optimal Cache Replacement for Graph AnalyticsVignesh Balaji, Neal Clayton Crago, Aamer Jaleel, Brandon LuciaHPCA 2021 · 被引用 41 次
- BlockGNN: Towards Efficient GNN Acceleration Using Block-Circulant Weight MatricesZhe Zhou, Bizhao Shi, Zhe Zhang, Yijin Guan 等DAC 2021 · 被引用 37 次
相关 Paper
- RTGA: A Redundancy-free Accelerator for High-Performance Temporal Graph Neural Network InferenceHui Yu, Yu Zhang, Andong Tan, Chenze Lu 等DAC 2024 · 被引用 5 次
- DiTile-DGNN: An Efficient Accelerator for Distributed Dynamic Graph Neural Network InferenceJiaqi Yang, Hao Zheng, Ahmed LouriISCA 2025 · 被引用 3 次
- Cambricon-DG: An Accelerator for Redundant-Free Dynamic Graph Neural Networks Based on Nonlinear IsolationZhifei Yue, Xinkai Song, Tianbo Liu, Xing Hu 等HPCA 2025 · 被引用 2 次
- TAGT: An Efficient Graph Transformer Accelerator with Topology-aware Sparsification and MergingHui Yu, Wei Zhang, Ligang He, Jin Zhao 等ISCA 2026
- I-DGNN: A Graph Dissimilarity-based Framework for Designing Scalable and Efficient DGNN AcceleratorsJiaqi Yang, Hao Zheng, Ahmed LouriHPCA 2025 · 被引用 3 次
