EOD: Enabling Low Latency GNN Inference via Near-Memory Concatenate Aggregation
Taehwan Kim, Yunki Han, Seohye Ha, Jiwan Kim, Lee-Sup Kim
Abstract
As online services based on graph databases increasingly integrate with machine learning, serving low-latency Graph Neural Network (GNN) inference for individual requests has become a critical challenge.Real-time GNN inference services operate in an inductive setup, which can handle newly added, previously unseen nodes and their edges.In this setup, the system must prepare the computational graph and input node features for target nodes, followed by GNN inference using the given input data.However, the workflow of a GNN serving system presents two key challenges that hinder low-latency inference.The first challenge arises from the extensive preparation step, which involves heavy memory access in host memory and significant data I/O to devices, constituting the largest portion of end-to-end inference latency.The second challenge is the well-known neighborhood explosion problem in GNN research.As the receptive field for target nodes increases exponentially with the number of layers, this issue exacerbates overall latency.To address these challenges, we propose a Near-Memory Processing (NMP) based low-latency GNN inference serving system named EOD.To ensure low-latency real-time GNN service, we co-design the algorithm and hardware to tackle the aforementioned issues.First, to mitigate the neighborhood explosion problem, we propose a precomputation method for the training node set, reducing memory access and computational complexity from exponential to linear growth.Additionally, we introduce a concatenated ZVC compression method to minimize the overhead of storing precomputed hidden features.Finally, to alleviate heavy host-side memory access and data I/O, we design an NMP architecture that enables efficient aggregation on concatenated ZVC-compressed data.As a result, EOD achieves a geometric mean of 981.1× and 912.0× aggregation speedup over the baseline and the existing architecture for GNN aggregation.Additionally, EOD achieves a geometric mean of 17.9×, and up to 74.3× end-to-end latency speedup over the GPU baseline.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get dafd0971-d91b-481c-9814-0a82473ee7a9Related papers
- FLAG: An FPGA-Based System for Low-Latency GNN Inference Service Using Vector QuantizationYunki Han, Taehwan Kim, Jiwan Kim, Seohye Ha et al.DAC 2025 · 1 citation
- D3-GNN: Dynamic Distributed Dataflow for Streaming Graph Neural NetworksRustam Guliyev, Aparajita Haldar, Hakan FerhatosmanogluVLDB 2024 · 5 citations
- LMC: Fast Training of GNNs via Subgraph Sampling with Provable ConvergenceZhihao Shi, Xize Liang, Jie WangICLR 2023 · 5 citations
- GraphFM: Improving Large-Scale GNN Training via Feature MomentumHaiyang Yu, Limei Wang, Bokun Wang, Meng Liu et al.ICML 2022 · 47 citations
- Serving Graph Compression for Graph Neural NetworksSi Si, Felix X. Yu, Ankit Singh Rawat, Cho-Jui Hsieh et al.ICLR 2023
