ScalaGraph: A Scalable Accelerator for Massively Parallel Graph Processing
Pengcheng Yao, Long Zheng, Yu Huang, Qinggang Wang, Chuangyi Gui, Zhen Zeng, Xiaofei Liao, Hai Jin, Jingling Xue
摘要
Graph processing is promising to extract valuable insights in graphs. Nowadays, emerging 3D-stacked memories and silicon technologies can provide over terabytes per second memory bandwidth and thousands of processing elements (PEs) to meet the high hardware demand of graph applications. However, this leap in hardware capability does not result in a huge increase but even a degradation sometimes in performance for graph processing. In this paper, we discover that the centralized on-chip memory hierarchy adopted in existing graph accelerators is the villain causing poor scalability due to its quadratic increase of hardware overheads with respect to the number of PEs.We present a novel distributed on-chip memory hierarchy by leveraging the network-on-chip (NoC) to enable massively parallel graph processing. We architect ScalaGraph, a brand new graph processing accelerator, to exploit this insight. ScalaGraph adopts a software-hardware co-design to minimize NoC communication overheads via an efficient row-oriented dataflow mapping and runtime aggregation. A specialized scheduling mechanism is also proposed to improve load imbalance. Our results on a Xilinx Alveo U280 FPGA card show that ScalaGraph on a modest configuration of 512 PEs achieves 2.2× and 3.2× speedups over a state-of-theart graph accelerator GraphDyns and a GPU-based graph system Gunrock, respectively. Moreover, ScalaGraph enables supporting at least 1,024 PEs with nearly linear performance scaling while GraphDyns fails to work.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- Pathfinding Future PIM Architectures by Demystifying a Commercial PIM TechnologyBongjoon Hyun, Taehun Kim, Dongjae Lee, Minsoo RhuHPCA 2024 · 被引用 62 次
- Piccolo: Large-Scale Graph Processing with Fine-Grained in-Memory Scatter-GatherChangmin Shin, Jaeyong Song, Hongsun Jang, Dogeun Kim 等HPCA 2025 · 被引用 5 次
- NOVA: A Novel Vertex Management Architecture for Scalable Graph ProcessingMarjan Fariborz, Mahyar Samani, Austin York, S. J. Ben Yoo 等HPCA 2025 · 被引用 2 次
- Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System SimulationYanjing Wang, Lizhou Wu, Sunfeng Gao, Yibo Tang 等HPCA 2026 · 被引用 1 次
- RidgeWalker: Perfectly Pipelined Graph Random Walks on FPGAsHongshi Tan, Yao Chen, Xinyu Chen, Qizhen Zhang 等HPCA 2026
相关 Paper
- Large-Scale Graph Processing on FPGAs with Caches for Thousands of Simultaneous MissesMikhail Asiatici, Paolo IenneISCA 2021 · 被引用 28 次
- A Locality-Aware Energy-Efficient Accelerator for Graph Mining ApplicationsPengcheng Yao, Long Zheng, Zhen Zeng, Yu Huang 等MICRO 2020 · 被引用 43 次
- OMeGa: Boosting Large-scale Graph Embeddings with Heterogeneous Memory ProcessingPeng Fang, Siqiang Luo, Fang Wang, Bolong Zheng 等ICDE 2025 · 被引用 1 次
- SCALE: A Structure-Centric Accelerator for Message Passing Graph Neural NetworksLingxiang Yin, Sanjay Gandham, Mingjie Lin, Hao ZhengMICRO 2024 · 被引用 4 次
- Hardware-Accelerated Hypergraph Processing with Chain-Driven SchedulingQinggang Wang, Long Zheng, Jingrui Yuan, Yu Huang 等HPCA 2022 · 被引用 9 次
