SG-Serve: Efficient Model Serving for Subgraph-based Graph Representation Learning
Qihui Zhou, Peiqi Yin, Xiao Yan, Changji Li, James Cheng
摘要
Subgraph-based graph representation learning (SGRL) is an emerging class of GNN models that achieve much higher accuracy for various tasks than classical GNN models (e.g., GCN and GAT). However, we observe that serving SGRL models for online applications is challenging due to their irregular request workloads . Specifically, some heavy requests need significantly more computation than regular requests (e.g., 100x), leading to excessively long tail latency in existing systems. As such, we build SG-Serve, which tailors subgraph extraction and model inference (i.e., the two main stages of SGRL models) to handle the irregular request workloads. For subgraph extraction on the CPU, we propose a general API to implement the diverse extraction methods of SGRL models. Beside generality, the API also exposes parallelization opportunities and allows the heavy requests to utilize multiple threads for speedup. To schedule the CPU threads to conduct extraction for concurrent requests, we design a work-stealing policy, which enjoys parallelism while avoiding head-of-line blocking. For model inference on the GPU, we batch the requests according to their workload instead of request count (i.e., as in existing systems) and run two GPU processes to prevent the heavy requests from monopolizing the GPU. Our experiments show that compared with existing systems, SG-Serve can reduce the 99th percentile (P99) latency by over 13x and improve the request throughput by over 33x.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- GENTI: GPU-powered Walk-based Subgraph Extraction for Scalable Representation Learning on Dynamic GraphsZihao Yu, Ningyi Liao, Siqiang LuoVLDB 2024 · 被引用 8 次
- Algorithm and System Co-design for Efficient Subgraph-based Graph Representation LearningHaoteng Yin, Muhan Zhang, Yanbang Wang, Jianguo Wang 等VLDB 2022 · 被引用 47 次
- SUREL+: Moving from Walks to Sets for Scalable Subgraph-based Graph Representation LearningHaoteng Yin, Muhan Zhang, Jianguo Wang, Pan LiVLDB 2023 · 被引用 13 次
- Translating Subgraphs to Nodes Makes Simple GNNs Strong and Efficient for Subgraph Representation LearningDongkwan Kim, Alice OhICML 2024 · 被引用 6 次
- BGL: GPU-Efficient GNN Training by Optimizing Graph Data I/O and PreprocessingTianfeng Liu, Yangrui Chen, Dan Li, Chuan Wu 等NSDI 2023
