Optimizing Inference Serving on Serverless Platforms
Ahsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia Smirni
摘要
Serverless computing is gaining popularity for machine learning (ML) serving workload due to its autonomous resource scaling, easy to use and pay-per-use cost model. Existing serverless platforms work well for image-based ML inference, where requests are homogeneous in service demands. That said, recent advances in natural language processing could not fully benefit from existing serverless platforms as their requests are intrinsically heterogeneous. Batching requests for processing can significantly increase ML serving efficiency while reducing monetary cost, thanks to the pay-per-use pricing model adopted by serverless platforms. Yet, batching heterogeneous ML requests leads to additional computation overhead as small requests need to be "padded" to the same size as large requests within the same batch. Reaching effective batching decisions (i.e., which requests should be batched together and why) is non-trivial: the padding overhead coupled with the serverless auto-scaling forms a complex optimization problem. To address this, we develop Multi-Buffer Serving (MBS), a framework that optimizes the batching of heterogeneous ML inference serving requests to minimize their monetary cost while meeting their service level objectives (SLOs). The core of MBS is a performance and cost estimator driven by analytical models supercharged by a Bayesian optimizer. MBS is prototyped and evaluated on AWS using bursty workloads. Experimental results show that MBS preserves SLOs while outperforming the state-of-the-art by up to 8 x in terms of cost savings while minimizing the padding overhead by up to 37 x with 3 x less number of serverless function invocations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- SpotServe: Serving Generative Large Language Models on Preemptible InstancesXupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi 等ASPLOS 2024 · 被引用 71 次
- HexGen: Generative Inference of Large Language Model over Heterogeneous EnvironmentYouhe Jiang, Ran Yan, Xiaozhe Yao, Yang Zhou 等ICML 2024 · 被引用 46 次
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo 等EuroSys 2024 · 被引用 29 次
- BlitzScale: Fast and Live Large Model Autoscaling with O(1) Host CachingDingyan Zhang, Haotian Wang, Yang Liu, Xingda Wei 等OSDI 2025 · 被引用 29 次
- FaaSGraph: Enabling Scalable, Efficient, and Cost-Effective Graph Processing with Serverless ComputingYushi Liu, Shixuan Sun, Zijun Li, Quan Chen 等ASPLOS 2024 · 被引用 14 次
它引用的顶会 Paper8
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 被引用 184 次
- SONIC: Application-aware Data Passing for Chained Serverless ApplicationsAshraf Mahgoub, Karthick Shankar, Subrata Mitra, Ana Klimovic 等USENIX ATC 2021 · 被引用 170 次
- AutoFocus: Efficient Multi-Scale InferenceMahyar Najibi, Bharat Singh, Larry DavisICCV 2019 · 被引用 143 次
- Lambada: Interactive Data Analytics on Cold Data Using Serverless Cloud InfrastructureIngo Müller, Renato Marroquín, Gustavo AlonsoSIGMOD 2020 · 被引用 135 次
- TurboTransformers: an efficient GPU serving system for transformer modelsJiarui Fang, Yang Yu, Chengduo Zhao, Jie ZhouPPoPP 2021 · 被引用 117 次
相关 Paper
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao 等HPCA 2026
- Optimizing Distributed Deployment of Mixture-of-Experts Model Inference in Serverless ComputingMengfan Liu, Wei Wang, Chuan WuINFOCOM 2025 · 被引用 6 次
- SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless ComputingChengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen 等SC 2024 · 被引用 8 次
- INFless: a native serverless system for low-latency, high-throughput inferenceYanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang 等ASPLOS 2022 · 被引用 145 次
- Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning InferenceYujeong Choi, Yunseong Kim, Minsoo RhuHPCA 2021 · 被引用 65 次
