TAGT: An Efficient Graph Transformer Accelerator with Topology-aware Sparsification and Merging
Hui Yu, Wei Zhang, Ligang He, Jin Zhao, Yu Zhang, Zixiao Wang
Abstract
Graph Transformers (GTs) have emerged as a powerful paradigm for graph representation learning, as their attention mechanism can capture long-range dependencies and model complex structural interactions beyond the local messagepassing scope of conventional Graph Neural Networks (GNNs). This capability has enabled GTs to achieve strong accuracy across important domains, including recommendation systems and VLSI congestion prediction. However, the global attention mechanism in GTs requires each vertex to attend to all other vertices, incurring O(N2) computation and intermediate data movement. As graph size increases, this quadratic complexity leads to prohibitive computational overhead and excessive offchip memory traffic, fundamentally limiting the scalability and efficiency of GT execution. In this paper, we propose TAGT, the first efficient topologyaware Graph Transformer accelerator designed to mitigate these performance bottlenecks. Specifically, we integrate a topologyaware sparsification and merging approach into the accelerator design that dramatically reduces the O(N2) complexity. TAGT introduces a structure-aware sparse subgraph, termed the Topology Dependency Subgraph (TDS), which exploits inherent topological dependencies and reduces the number of attended edges to O(N log N) on average. The TDS is designed to retain local neighborhood structure while capturing essential higherorder interrelationships. By performing attention on the TDS, TAGT approximates global attention over the entire graph with negligible accuracy loss while eliminating most unnecessary computations and off-chip data movements. To fully harness the performance potential of this approach, TAGT incorporates a datadriven loading and merging engine to minimize off-chip memory accesses and reduce TDS construction overhead on the fly. TAGT also introduces a TDS-based fast attention unit to improve the parallelism of attention computation. We implement and evaluate TAGT on a Xilinx Alveo U280 FPGA card. Experimental results show that TAGT achieves average speedups of 175.4 × and 18.6 ×, together with energy savings of 217.2 × and 24.8 ×, over state-ofthe-art software GT solutions on Intel Xeon CPUs and NVIDIA A100 GPUs, respectively. Compared with representative GNN accelerators, including FlowGNN, MEGA, and BingoGCN, TAGT delivers average speedups of 8.2 ×, 6.9 ×, and 4.7 ×, and energy savings of 9.3 ×, 7.5 ×, and 5.2 ×, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 5a610362-4ef5-4047-8417-a5f2eabd5b1eRelated papers
- TaGNN: An Efficient Topology-aware Accelerator for High-performance Dynamic Graph Neural NetworkHui Yu, Yu Zhang, Ligang He, Bing Peng et al.SC 2025 · 2 citations
- RTGA: A Redundancy-free Accelerator for High-Performance Temporal Graph Neural Network InferenceHui Yu, Yu Zhang, Andong Tan, Chenze Lu et al.DAC 2024 · 5 citations
- A Scalable and Effective Alternative to Graph TransformersKaan Sancak, Zhigang Hua, Jin Fang, Yan Xie et al.AAAI 2025 · 5 citations
- DUALFormer: Dual Graph TransformerJiaming Zhuo, Yuwei Liu, Yintong Lu, Ziyi Ma et al.ICLR 2025
- RAHP: A Redundancy-aware Accelerator for High-performance Hypergraph Neural NetworkHui Yu, Yu Zhang, Ligang He, Yingqi Zhao et al.MICRO 2024 · 6 citations
