ScalaGraph: A Scalable Accelerator for Massively Parallel Graph Processing
Pengcheng Yao, Long Zheng, Yu Huang, Qinggang Wang, Chuangyi Gui, Zhen Zeng, Xiaofei Liao, Hai Jin, Jingling Xue
Abstract
Graph processing is promising to extract valuable insights in graphs. Nowadays, emerging 3D-stacked memories and silicon technologies can provide over terabytes per second memory bandwidth and thousands of processing elements (PEs) to meet the high hardware demand of graph applications. However, this leap in hardware capability does not result in a huge increase but even a degradation sometimes in performance for graph processing. In this paper, we discover that the centralized on-chip memory hierarchy adopted in existing graph accelerators is the villain causing poor scalability due to its quadratic increase of hardware overheads with respect to the number of PEs.We present a novel distributed on-chip memory hierarchy by leveraging the network-on-chip (NoC) to enable massively parallel graph processing. We architect ScalaGraph, a brand new graph processing accelerator, to exploit this insight. ScalaGraph adopts a software-hardware co-design to minimize NoC communication overheads via an efficient row-oriented dataflow mapping and runtime aggregation. A specialized scheduling mechanism is also proposed to improve load imbalance. Our results on a Xilinx Alveo U280 FPGA card show that ScalaGraph on a modest configuration of 512 PEs achieves 2.2× and 3.2× speedups over a state-of-theart graph accelerator GraphDyns and a GPU-based graph system Gunrock, respectively. Moreover, ScalaGraph enables supporting at least 1,024 PEs with nearly linear performance scaling while GraphDyns fails to work.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 0db23568-5dd4-4b84-a449-0a36682c7529Cited by top-tier papers6
- Pathfinding Future PIM Architectures by Demystifying a Commercial PIM TechnologyBongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo RhuHPCA 2024 · 62 citations
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim et al.HPCA 2025 · 5 citations
- NOVA: A Novel Vertex Management Architecture for Scalable Graph ProcessingMarjan Fariborz, Mahyar Samani, Austin York, S. J. Ben Yoo et al.HPCA 2025 · 2 citations
- Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System SimulationYanjing Wang, Lizhou Wu, Sunfeng Gao, Yibo Tang et al.HPCA 2026 · 1 citation
- RidgeWalker: Perfectly Pipelined Graph Random Walks on FPGAsHongshi Tan, Yao Chen, Xinyu Chen, Qizhen Zhang et al.HPCA 2026
Related papers
- Large-Scale Graph Processing on FPGAs with Caches for Thousands of Simultaneous MissesMikhail Asiatici, Paolo IenneISCA 2021 · 28 citations
- A Locality-Aware Energy-Efficient Accelerator for Graph Mining ApplicationsPengcheng Yao, Long Zheng, Zhen Zeng, Yu Huang et al.MICRO 2020 · 43 citations
- OMeGa: Boosting Large-scale Graph Embeddings with Heterogeneous Memory ProcessingPeng Fang, Siqiang Luo, Fang Wang, Bolong Zheng et al.ICDE 2025 · 1 citation
- SCALE: A Structure-Centric Accelerator for Message Passing Graph Neural NetworksLingxiang Yin, Sanjay Gandham, Mingjie Lin, Hao ZhengMICRO 2024 · 4 citations
- Hardware-Accelerated Hypergraph Processing with Chain-Driven SchedulingQinggang Wang, Long Zheng, Jingrui Yuan, Yu Huang et al.HPCA 2022 · 9 citations
