λGrapher: A Resource-Efficient Serverless System for GNN Serving through Graph Sharing
Haichuan Hu, Fangming Liu, Qiangyu Pei, Yongjie Yuan, Zichen Xu, Lin Wang
摘要
Graph Neural Networks (GNNs) have been increasingly adopted for graph analysis in web applications such as social networks. Yet, efficient GNN serving remains a critical challenge due to high workload fluctuations and intricate GNN operations. Serverless computing, thanks to its flexibility and agility, offers on-demand serving of GNN inference requests. Alas, the request-centric serverless model is still too coarse-grained to avoid resource waste. Observing the significant data locality in computation graphs of requests, we propose λGrapher, a serverless system for GNN serving that achieves resource efficiency through graph sharing and fine-grained resource allocation. "Grapher features the following designs: (1) adaptive timeout for request buffering to balance resource efficiency and inference latency, (2) graph-centric scheduling to minimize computation and memory redundancy, and (3) resource-centric function management with fine-grained resource allocation catered to the resource sensitivities of GNN operations and function orchestration optimized to hide communication latency. We implement a prototype of λGrapher based on the representative open-source serverless platform Knative and evaluate it with real-world traces from various web applications. Our results show that λGrapher can achieve an average savings of 61.5% in memory resource and 47.2% in computing resource compared with the state of the arts while ensuring GNN inference latency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- An LLM-Guided Query-Aware Inference System for GNN Models on Large Knowledge GraphsWaleed Afandi, Hussein Abdallah, Ashraf Aboulnaga, Essam MansourICDE 2026
- Incremental GNN Embedding Computation on Streaming GraphsQiange Wang, Haoran Lv, Yanfeng Zhang, Weng-Fai Wong 等ICDE 2026
它引用的顶会 Paper12
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu 等OSDI 2023 · 被引用 211 次
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless ThreadsJohn Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng 等OSDI 2021 · 被引用 175 次
- GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUsYuke Wang, Boyuan Feng, Gushu Li, Shuangchen Li 等OSDI 2021 · 被引用 163 次
- Hardware Acceleration of Graph Neural NetworksAdam Auten, Matthew Tomei, Rakesh KumarDAC 2020 · 被引用 108 次
- ByteGNN: Efficient Graph Neural Network Training at Large ScaleChenguang Zheng, Hongzhi Chen, Yuxuan Cheng, Zhezheng Song 等VLDB 2022 · 被引用 107 次
相关 Paper
- FaaSGraph: Enabling Scalable, Efficient, and Cost-Effective Graph Processing with Serverless ComputingYushi Liu, Shixuan Sun, Zijun Li, Quan Chen 等ASPLOS 2024 · 被引用 14 次
- FaaSBoard: Efficient Graph Processing with a Disaggregated Architecture on Serverless ServicesYushi Liu, Yikang Ruan, Letian Ruan, Zijun Li 等SIGMOD 2026
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao 等HPCA 2026
- Rocket: Warming Serverless Inference via Hierarchical ML Artifact Pre-loading and SharingXiaofei Yue, Song Yang, Fan Li, Youqi Li 等INFOCOM 2026 · 被引用 2 次
- StreamBox: A Lightweight GPU SandBox for Serverless Inference WorkflowHao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim 等USENIX ATC 2024 · 被引用 21 次
