λGrapher: A Resource-Efficient Serverless System for GNN Serving through Graph Sharing
Haichuan Hu, Fangming Liu, Qiangyu Pei, Yongjie Yuan, Zichen Xu, Lin Wang
Abstract
Graph Neural Networks (GNNs) have been increasingly adopted for graph analysis in web applications such as social networks. Yet, efficient GNN serving remains a critical challenge due to high workload fluctuations and intricate GNN operations. Serverless computing, thanks to its flexibility and agility, offers on-demand serving of GNN inference requests. Alas, the request-centric serverless model is still too coarse-grained to avoid resource waste. Observing the significant data locality in computation graphs of requests, we propose λGrapher, a serverless system for GNN serving that achieves resource efficiency through graph sharing and fine-grained resource allocation. "Grapher features the following designs: (1) adaptive timeout for request buffering to balance resource efficiency and inference latency, (2) graph-centric scheduling to minimize computation and memory redundancy, and (3) resource-centric function management with fine-grained resource allocation catered to the resource sensitivities of GNN operations and function orchestration optimized to hide communication latency. We implement a prototype of λGrapher based on the representative open-source serverless platform Knative and evaluate it with real-world traces from various web applications. Our results show that λGrapher can achieve an average savings of 61.5% in memory resource and 47.2% in computing resource compared with the state of the arts while ensuring GNN inference latency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 570dd92e-1676-4d37-bc4f-1b30e4f67fa5Cited by top-tier papers2
- An LLM-Guided Query-Aware Inference System for GNN Models on Large Knowledge GraphsWaleed Afandi, Hussein Abdallah, Ashraf Aboulnaga, Essam MansourICDE 2026
- Incremental GNN Embedding Computation on Streaming GraphsQiange Wang, Haoran Lv, Yanfeng Zhang, Weng-Fai Wong et al.ICDE 2026
Builds on12
- AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning ServingZhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu et al.OSDI 2023 · 211 citations
- Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless ThreadsJohn Thorpe, Yifan Qiao, Jonathan Eyolfson, Shen Teng et al.OSDI 2021 · 175 citations
- GNNAdvisor: An Adaptive and Efficient Runtime System for GNN Acceleration on GPUsYuke Wang, Boyuan Feng, Gushu Li, Shuangchen Li et al.OSDI 2021 · 163 citations
- Hardware Acceleration of Graph Neural NetworksAdam Auten, Matthew Tomei, Rakesh KumarDAC 2020 · 108 citations
- ByteGNN: Efficient Graph Neural Network Training at Large ScaleChenguang Zheng, Hongzhi Chen, Yuxuan Cheng, Zhezheng Song et al.VLDB 2022 · 107 citations
Related papers
- FaaSGraph: Enabling Scalable, Efficient, and Cost-Effective Graph Processing with Serverless ComputingYushi Liu, Shixuan Sun, Zijun Li, Quan Chen et al.ASPLOS 2024 · 14 citations
- FaaSBoard: Efficient Graph Processing with a Disaggregated Architecture on Serverless ServicesYushi Liu, Yikang Ruan, Letian Ruan, Zijun Li et al.SIGMOD 2026
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao et al.HPCA 2026
- Rocket: Warming Serverless Inference via Hierarchical ML Artifact Pre-loading and SharingXiaofei Yue, Song Yang, Fan Li, Youqi Li et al.INFOCOM 2026 · 2 citations
- StreamBox: A Lightweight GPU SandBox for Serverless Inference WorkflowHao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim et al.USENIX ATC 2024 · 21 citations
