Clementi: Efficient Load Balancing and Communication Overlap for Multi-FPGA Graph Processing
Feng Yu, Hongshi Tan, Xinyu Chen, Yao Chen, Bingsheng He, Weng-Fai Wong
摘要
Efficient graph processing is critical in various modern applications, such as social network analysis, recommendation systems, and large-scale data mining. Traditional single-FPGA systems struggle to handle the increasing size and complexity of real-world graphs due to limitations in memory and computational resources. Existing multi-FPGA solutions face significant challenges, including high communication overhead caused by irregular data transfer patterns and workload imbalances stemming from skewed graph distributions. These inefficiencies hinder scalability and performance, highlighting a critical research gap. To address these issues, we introduce Clementi, an efficient multi-FPGA graph processing framework that features customized fine-grained pipelines for computation and cross-FPGA communication. Clementi uniquely integrates an accurate performance model for execution time prediction, enabling a novel scheduling method that balances workload distribution and minimizes communication overhead by overlapping communication and computation stages. Experimental results demonstrate that Clementi achieves speedups of up to 8.75× compared to state-of-the-art multi-FPGA designs, indicating significant improvements in processing efficiency as the number of FPGAs increases. This near-linear scalability underscores the framework' s potential to enhance graph processing capabilities in practical applications. Clementi is open-sourced at https://github.com/Xtra-Computing/Clementi.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- ReGraph: Scaling Graph Processing on HBM-enabled FPGAs with Heterogeneous PipelinesXinyu Chen, Yao Chen, Feng Cheng, Hongshi Tan 等MICRO 2022 · 被引用 46 次
- OneShotSTL: One-Shot Seasonal-Trend Decomposition For Online Time Series Anomaly Detection And ForecastingXiao He, Ye Li, Jian Tan, Bin Wu 等VLDB 2023 · 被引用 40 次
- Large-Scale Graph Processing on FPGAs with Caches for Thousands of Simultaneous MissesMikhail Asiatici, Paolo IenneISCA 2021 · 被引用 28 次
相关 Paper
- MGG: Accelerating Graph Neural Networks with Fine-Grained Intra-Kernel Communication-Computation Pipelining on Multi-GPU PlatformsYuke Wang, Boyuan Feng, Zheng Wang, Tong Geng 等OSDI 2023 · 被引用 46 次
- CGgraph: An Ultra-fast Graph Processing System on Modern Commodity CPU-GPU Co-processorPengjie Cui, Haotian Liu, Bo Tang, Ye YuanVLDB 2024 · 被引用 18 次
- HyTGraph: GPU-Accelerated Graph Processing with Hybrid Transfer ManagementQiange Wang, Xin Ai, Yanfeng Zhang, Jing Chen 等ICDE 2023 · 被引用 14 次
- SympleGraph: distributed graph processing with precise loop-carried dependency guaranteeYouwei Zhuo, Jingji Chen, Qinyi Luo, Yanzhi Wang 等PLDI 2020 · 被引用 13 次
- Cache-Efficient Fork-Processing Patterns on Large GraphsShengliang Lu, Shixuan Sun, Johns Paul, Yuchen Li 等SIGMOD 2021 · 被引用 10 次
