Towards Resource-Efficient Serverless LLM Inference with SLINFER
Chuhao Xu, Zijun Li, Quan Chen, Han Zhao, Xueyan Tang, Minyi Guo
Abstract
The rise of LLMs has driven demand for private serverless deployments, characterized by moderate-sized models and infrequent requests. While existing serverless solutions follow exclusive GPU allocation, we take a step back to explore modern platforms and find that: Emerging CPU architectures with built-in accelerators are capable of serving LLMs but remain underutilized, and both CPUs and GPUs can accommodate multiple LLMs simultaneously. We propose SLINFER, a resource-efficient serverless inference scheme tailored for small- to mid-sized LLMs that enables elastic and on-demand sharing across heterogeneous hardware. SLINFER tackles three fundamental challenges: (1) precise, fine-grained compute resource allocation at token-level to handle fluctuating computational demands; (2) a coordinated and forward-looking memory scaling mechanism to detect out-ofmemory hazards and reduce operational overhead; and (3) a dual approach that consolidates fragmented instances through proactive preemption and reactive bin-packing. Experimental results on 4 32-core CPUs and 4 A100 GPUs show that SLINFER improves serving capacity by 47% - 62% through sharing, while further leveraging CPUs boosts this to 86% - 154%.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f7f99b97-4033-49d0-b053-33c664996bddBuilds on38
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
Related papers
- PARIS and ELSA: an elastic scheduling algorithm for reconfigurable multi-GPU inference serversYunseong Kim, Yujeong Choi, Minsoo RhuDAC 2022 · 15 citations
- PiLLM: Resource-Efficient LLM Inference Using Workload PredictionYunqian Fan, Shihao Bai, Ruihao Gong, Zaijun Wang et al.EuroSys 2026
- Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient InferenceMinchen Yu, Ao Wang, Dong Chen, Haoxuan Yu et al.USENIX ATC 2025
- Dilu: Enabling GPU Resourcing-on-Demand for Serverless DL Serving via Introspective ElasticityCunchi Lv, Xiao Shi, Zhengyu Lei, Jinyue Huang et al.ASPLOS 2025 · 10 citations
- Optimizing Inference Serving on Serverless PlatformsAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniVLDB 2022 · 76 citations
