SkyServe: Serving AI Models across Regions and Clouds with Spot Instances
Ziming Mao, Tian Xia, Zhanghao Wu, Wei-Lin Chiang, Tyler Griggs, Romil Bhardwaj, Zongheng Yang, Scott Shenker, Ion Stoica
摘要
Recent years have witnessed an explosive growth of AI models. The high cost of hosting AI services on GPUs and their demanding service requirements, make it timely and challenging to lower service costs and guarantee service quality. While spot instances have long been offered with a large discount, spot preemptions have discouraged users from using them to host model replicas when serving AI models.
To address this, we propose a simple yet efficient policy, SpotHedge, that leverages spot replicas across different failure domains (e.g., regions and clouds) to ensure availability, lower costs, and high service quality. SpotHedge intelligently spreads spot replicas across different regions and clouds to improve availability and reduce correlated preemptions, overprovisions cheap spot replicas than required as a safeguard against possible preemptions, and dynamically falls back to on-demand replicas when spot replicas become unavailable. We built SkyServe, a system leveraging SpotHedge to efficiently serve AI models over a mixture of spot and on-demand replicas across regions and clouds. We compared SkyServe with both research and production systems on real AI workloads: SkyServe reduces cost by 43% on average while achieving high resource availability compared to using on-demand replicas. Additionally, SkyServe improves P50, P90, and P99 latency by 2.3×, 2.1×, 2.1× on average compared to other research and production systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- RLBoost: Harvesting Preemptible Cloud Resources for Cost-Efficient Reinforcement Learning on LLMsYongji Wu, Xueshen Liu, Haizhong Zheng, Juncheng Gu 等NSDI 2026 · 被引用 4 次
- SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM InferenceTian Xia, Ziming Mao, Jamison Kerney, Ethan J. Jackson 等EuroSys 2026 · 被引用 2 次
- Sailor: Automating Distributed Training over Dynamic, Heterogeneous, and Geo-distributed ClustersFoteini Strati, Zhendong Zhang, George Manos, Ixeia Sánchez Périz 等SOSP 2025 · 被引用 2 次
- Efficient Remote KV Cache Reuse with GPU-native Video CodecLiang Mi, Weijun Wang, Jinghan Chen, Ting Cao 等SIGCOMM 2026
- Prompt Inversion Attack Against Collaborative Inference of Large Language ModelsWenjie Qu, Yuguang Zhou, Yongji Wu, Tingsong Xiao 等S&P 2025
它引用的顶会 Paper13
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
相关 Paper
- SpotServe: Serving Generative Large Language Models on Preemptible InstancesXupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi 等ASPLOS 2024 · 被引用 71 次
- Can't Be Late: Optimizing Spot Instance Savings under DeadlinesZhanghao Wu, Wei-Lin Chiang, Ziming Mao, Zongheng Yang 等NSDI 2024 · 被引用 40 次
- Snape: Reliable and Low-Cost Computing with Mixture of Spot and On-Demand VMsFangkai Yang, Lu Wang, Zhenyu Xu, Jue Zhang 等ASPLOS 2023 · 被引用 18 次
- GFS: A Preemption-aware Scheduling Framework for GPU Clusters with Predictive Spot Instance ManagementJiaang Duan, Shenglin Xu, Shiyou Qian, Dingyu Yang 等ASPLOS 2026 · 被引用 1 次
- Cremes: Cost-Efficient and Reliable Microservice Execution on Spot InstancesLiao Chen, Chenyu Lin, Junlin Chen, Shutian Luo 等HPDC 2026
