ReGraph: Scaling Graph Processing on HBM-enabled FPGAs with Heterogeneous Pipelines
Xinyu Chen, Yao Chen, Feng Cheng, Hongshi Tan, Bingsheng He, Weng-Fai Wong
摘要
The use of FPGAs for efficient graph processing has attracted significant interest. Recent memory subsystem upgrades including the introduction of HBM in FPGAs promise to further alleviate memory bottlenecks. However, modern multi-channel HBM requires much more processing pipelines to fully utilize its bandwidth potential. Existing designs do not scale well, resulting in underutilization of the HBM facilities even when all other resources are fully consumed.
In this paper, we re-examined the graph processing workloads and found much diversity in processing. We also found that the diverse workloads can be easily classified into two types, namely dense and sparse partitions. This motivates us to propose a resource-efficient heterogeneous pipeline architecture. Our heterogeneous architecture comprises of two types of pipelines: Little pipelines to process dense partitions with good locality and Big pipelines to process sparse partitions with extremely poor locality. Unlike traditional monolithic pipeline designs, the heterogeneous pipelines are tailored for more specific memory access patterns, and hence are more lightweight, allowing the architecture to scale up more effectively with limited resources. In addition, we propose a model-guided task scheduling method that schedules partitions to the right pipeline types, generates the most efficient pipeline combination and balances workloads. Furthermore, we develop an automated open-source framework, called ReGraph 1 , which automates the entire development process. ReGraph outperforms state-of-the-art FPGA accelerators by up to 5.9× in terms of performance and 12× in terms of resource efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Ecco: Improving Memory Bandwidth and Capacity for LLMs via Entropy-Aware Cache CompressionFeng Cheng, Cong Guo, Chiyue Wei, Junyao Zhang 等ISCA 2025 · 被引用 12 次
- DECA: A Near-Core LLM Decompression Accelerator Grounded on a 3D Roofline ModelGerasimos Gerogiannis, Stijn Eyerman, Evangelos Georganas, Wim Heirman 等MICRO 2025 · 被引用 5 次
- Clementi: Efficient Load Balancing and Communication Overlap for Multi-FPGA Graph ProcessingFeng Yu, Hongshi Tan, Xinyu Chen, Yao Chen 等SIGMOD 2025 · 被引用 3 次
- HiT: A Unified Sparsity-Adaptive Architecture for High-Throughput Matrix MultiplicationTingting Xiang, Xiaochen Wang, Miao Yu, Trevor E. CarlsonISCA 2026
- RidgeWalker: Perfectly Pipelined Graph Random Walks on FPGAsHongshi Tan, Yao Chen, Xinyu Chen, Qizhen Zhang 等HPCA 2026
它引用的顶会 Paper2
相关 Paper
- FALA: Locality-Aware PIM-Host Cooperation for Graph Processing with Fine-Grained Column AccessChangmin Shin, Jaeyong Song, Seongmin Na, Jun Sung 等MICRO 2025 · 被引用 5 次
- PolyGraph: Exposing the Value of Flexibility for Graph Processing AcceleratorsVidushi Dadu, Sihao Liu, Tony NowatzkiISCA 2021 · 被引用 60 次
- HedraRAG: Co-Optimizing Generation and Retrieval for Heterogeneous RAG WorkflowsZhengding Hu, Vibha Murthy, Zaifeng Pan, Wanlu Li 等SOSP 2025 · 被引用 1 次
- Understand and Accelerate Memory Processing Pipeline for Large Language Model InferenceZifan He, Rui Ma, Yizhou Sun, Jason CongICML 2026
- SparseWeaver: Converting Sparse Operations as Dense Operations on GPUs for Graph WorkloadsShinnung Jeong, Liam Paul Cooper, Ju Min Lee, Heelim Choi 等HPCA 2025 · 被引用 2 次
