Astraea: towards QoS-aware and resource-efficient multi-stage GPU services
Wei Zhang, Quan Chen, Kaihua Fu, Ningxin Zheng, Zhiyi Huang, Jingwen Leng, Minyi Guo
摘要
Multi-stage user-facing applications on GPUs are widely-used nowa- days, and are often implemented to be microservices. Prior re- search works are not applicable to ensuring QoS of GPU-based microservices due to the different communication patterns and shared resource contentions. We propose Astraea to manage GPU microservices considering the above factors. In Astraea, a microser- vice deployment policy is used to maximize the supported peak service load while ensuring the required QoS. To adaptively switch the communication methods between microservices according to different deployments, we propose an auto-scaling GPU communi- cation framework. The framework automatically scales based on the currently used hardware topology and microservice location, and adopts global memory-based techniques to reduce intra-GPU communication. Astraea increases the supported peak load by up to 82.3% while achieving the desired 99%-ile latency target compared with state-of-the-art solutions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper6
- Nodens: Enabling Resource Efficient and Fast QoS Recovery of Dynamic Microservice Applications in DatacentersJiuchen Shi, Hang Zhang, Zhixin Tong, Quan Chen 等USENIX ATC 2023 · 被引用 32 次
- Low-Overhead General-Purpose Near-Data Processing in CXL Memory ExpandersHyungkyu Ham, Jeongmin Hong, Geonwoo Park, Yunseon Shin 等MICRO 2024 · 被引用 26 次
- StreamBox: A Lightweight GPU SandBox for Serverless Inference WorkflowHao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim 等USENIX ATC 2024 · 被引用 21 次
- Multiplexing Dynamic Deep Learning Workloads with SLO-awareness in GPU ClustersWenyan Chen, Chengzhi Lu, Huanle Xu, Kejiang Ye 等EuroSys 2025 · 被引用 13 次
- SMIless: Serving DAG-based Inference with Dynamic Invocations under Serverless ComputingChengzhi Lu, Huanle Xu, Yudan Li, Wenyan Chen 等SC 2024 · 被引用 8 次
相关 Paper
- Erlang: Application-Aware Autoscaling for Cloud MicroservicesVighnesh Sachidananda, Anirudh SivaramanEuroSys 2024 · 被引用 7 次
- Enable simultaneous DNN services based on deterministic operator overlap and precise latency predictionWeihao Cui, Han Zhao, Quan Chen, Ningxin Zheng 等SC 2021 · 被引用 62 次
- gShare: Efficient GPU Sharing with Aggressive Scheduling in Multi-tenant FaaS platformYanan Yang, Zhengxiong Jiang, Meiqi Zhu, Hongqiang Xu 等ASPLOS 2026 · 被引用 1 次
- Automatic Policy Generation for Inter-Service Access Control of MicroservicesXing Li, Yan Chen, Zhiqiang Lin, Xiao Wang 等USENIX Security 2021 · 被引用 64 次
- Autothrottle: A Practical Bi-Level Approach to Resource Management for SLO-Targeted MicroservicesZibo Wang, Pinghe Li, Chieh-Jan Mike Liang, Feng Wu 等NSDI 2024
