DEEPSERVE: Serverless Large Language Model Serving at Scale
Junhao Hu, Jiang Xu, Zhixia Liu, Yulong He, Yuetao Chen, Hao Xu, Jiang Liu, Jie Meng, Baoquan Zhang, Shining Wan, Gengyuan Dan, Zhiyu Dong
Abstract
In this paper, we propose DEEPSERVE, a scalable and serverless AI platform designed to efficiently serve large language models (LLMs) at scale in cloud environments. DEEPSERVE addresses key challenges such as resource allocation, serving efficiency, and cold start latencies through four main design components. First, DEEPSERVE uses a simple serverless abstraction called the request-job-task model, which helps manage diverse AI workloads across posttraining and model-serving tasks. Second, DEEPSERVE integrates an in-house serving engine named FLOWSERVE using a microkernel-inspired design, NPU-centric execution, and SPMD-based parallelism to optimize LLM serving. Third, DEEPSERVE includes novel scheduling policies tailored for a configuration with both PD-disaggregated and PD-colocated instances. Fourth, DEEPSERVE includes optimizations such as pre-warmed pods, DRAM pre-loading, and NPU-fork, which allow DEEPSERVE to scale up to 64 instances in seconds. DEEPSERVE has been in production for over a year, operating on a large Ascend NPU cluster and providing industrystandard APIs for fine-tuning, agent serving, and model serving to our customers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 39e7f8f5-82e9-4e18-a554-600e1ef139f7Cited by top-tier papers4
- HydraServe: Minimizing Cold Start Latency for Serverless LLM Serving in Public CloudsChiheng Lou, Sheng Qi, Chao Jin, Dapeng Nie et al.NSDI 2026 · 22 citations
- CacheSlide: Unlocking Cross Position-Aware KV Cache Reuse for Accelerating LLM ServingYang Liu, Yunfei Gu, Liqiang Zhang, Chentao Wu et al.FAST 2026 · 14 citations
- Accelerating Model Loading in LLM Inference by Programmable Page CacheYubo Liu, Hongbo Li, Xiaojia Huang, Yongfeng Wang et al.FAST 2026
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao et al.HPCA 2026
Builds on19
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- Mooncake: Trading More Storage for Less Computation - A KVCache-centric Architecture for Serving LLM ChatbotRuoyu Qin, Zheming Li, Weiran He, Jialei Cui et al.FAST 2025 · 337 citations
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah et al.ISCA 2024 · 282 citations
Related papers
- XY-Serve: End-to-End Versatile Production Serving for Dynamic LLM WorkloadsMingcong Song, Xinru Tang, Fengfan Hou, Jing Li et al.ASPLOS 2026
- WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM ServingChiheng Lou, Sheng Qi, Rui Kang, Yong Zhang et al.ICML 2026 · 3 citations
- SpotServe: Serving Generative Large Language Models on Preemptible InstancesXupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi et al.ASPLOS 2024 · 71 citations
- FastServe: Iteration-Level Preemptive Scheduling for Large Language Model InferenceBingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu et al.NSDI 2026 · 12 citations
- WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic SchedulingJingqi Feng, Yukai Huang, Rui Zhang, Sicheng Liang et al.ISCA 2025 · 16 citations
