Astraea: towards QoS-aware and resource-efficient multi-stage GPU services
Wei Zhang, Quan Chen, Kaihua Fu, Ningxin Zheng, Zhiyi Huang, Jingwen Leng, Minyi Guo
Abstract
Multi-stage user-facing applications on GPUs are widely-used nowa- days, and are often implemented to be microservices. Prior re- search works are not applicable to ensuring QoS of GPU-based microservices due to the different communication patterns and shared resource contentions. We propose Astraea to manage GPU microservices considering the above factors. In Astraea, a microser- vice deployment policy is used to maximize the supported peak service load while ensuring the required QoS. To adaptively switch the communication methods between microservices according to different deployments, we propose an auto-scaling GPU communi- cation framework. The framework automatically scales based on the currently used hardware topology and microservice location, and adopts global memory-based techniques to reduce intra-GPU communication. Astraea increases the supported peak load by up to 82.3% while achieving the desired 99%-ile latency target compared with state-of-the-art solutions.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 140c164a-fa7b-4560-bfe7-343fc9ac5aa4Cited by top-tier papers6
- Nodens: Enabling Resource Efficient and Fast QoS Recovery of Dynamic Microservice Applications in DatacentersJiuchen Shi, Hang Zhang, Zhixin Tong, Quan Chen et al.USENIX ATC 2023 · 32 citations
- Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersHyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin et al.MICRO 2024 · 26 citations
- StreamBox: A Lightweight GPU SandBox for Serverless Inference WorkflowHao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim et al.USENIX ATC 2024 · 21 citations
- Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye et al.EuroSys 2025 · 13 citations
- SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless ComputingChengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen et al.SC 2024 · 8 citations
Related papers
- Erlang: Application-Aware Autoscaling for Cloud MicroservicesVighnesh Sachidananda, Anirudh SivaramanEuroSys 2024 · 7 citations
- Enable simultaneous DNN services based on deterministic operator overlap and precise latency predictionWeihao Cui, Han Zhao, Quan Chen, Ningxin Zheng et al.SC 2021 · 62 citations
- gShare: Efficient GPU Sharing with Aggressive Scheduling in Multi-tenant FaaS platformYanan Yang, Zhengxiong Jiang, Meiqi Zhu, Hongqiang Xu et al.ASPLOS 2026 · 1 citation
- Automatic Policy Generation for Inter-Service Access Control of MicroservicesXing Li, Yan Chen, Zhiqiang Lin, Xiao Wang et al.USENIX Security 2021 · 64 citations
- Autothrottle: A Practical Bi-Level Approach to Resource Management for SLO-Targeted MicroservicesZibo Wang, Pinghe Li, Chieh-Jan Mike Liang, Feng Wu et al.NSDI 2024
