Graphite: A NUMA-aware HPC System for Graph Analytics Based on a new MPI * X Parallelism Model
Mohammad Hasanzadeh-Mofrad, Rami G. Melhem, Muhammad Yousuf Ahmad, Mohammad Hammoud
Abstract
In this paper, we propose a new parallelism model denoted as MPI * X and suggest a linear algebra-based graph analytics system, namely, Graphite, which effectively employs it. MPI * X promotes thread-based partitioning to distribute computation and communication across threads on a cluster of machines, while eliminating the need for unnecessary thread synchronizations. Consequently, it contrasts with the traditional MPI + X parallelism model, which utilizes process-based partitioning to distribute data among processes as a way to scale out on a cluster of machines (the MPI part), then splits each partition into subpartitions among the threads of each process as a method to scale up within a machine (the X part). Besides adopting MPI * X, Graphite is NUMA-aware. In particular, it assigns threads to partitions in a way that exploits CPU and memory affinity, alongside leveraging faster MPI shared memory transport. Moreover, it adopts a variant of the popular GAS (Gather, Apply, and Scatter) computing model, thus decoupling the computation of partitions from the communication of partial results. Lastly, it supports thread-level asynchrony, which does not only overlap the computation with communication, but further interleaves multiple communications. We compared Graphite against GraphPad, Gemini, and LA3 graph analytics systems in an HPC environment using different graph applications. Results show that Graphite is roughly up to 3X faster than these state-of-the-art systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45130243-b600-48ad-a02c-abfc63f1c078Cited by top-tier papers2
- FaaSGraph: Enabling Scalable, Efficient, and Cost-Effective Graph Processing with Serverless ComputingYushi Liu, Shixuan Sun, Zijun Li, Quan Chen et al.ASPLOS 2024 · 14 citations
- Pluto: High-Performance, Memory-Efficient Distributed Graph Analytics through Advanced MirroringYing-Wei Wu, Christopher J. Rossbach, Mattan ErezOSDI 2026
Related papers
- GX-Plug: a Middleware for Plugging Accelerators to Distributed Graph ProcessingKai Zou, Xike Xie, Qi Li, Deyu KongICDE 2022 · 2 citations
- Graph Computation with Adaptive GranularityRuiqi Xu, Yue Wang, Xiaokui XiaoICDE 2024
- GraphCube: Interconnection Hierarchy-aware Graph ProcessingXinbiao Gan, Guang Wu, Shenghao Qiu, Feng Xiong et al.PPoPP 2024 · 15 citations
- Cache-Efficient Fork-Processing Patterns on Large GraphsShengliang Lu, Shixuan Sun, Johns Paul, Yuchen Li et al.SIGMOD 2021 · 10 citations
- TianheEngine: Hierarchy-aware Adaptive Partitioning System for Trillion-scale Graph ProcessingXinbiao Gan, Tiejun Li, Yiqi Wang, Qiang Zhang et al.SC 2025 · 1 citation
