Automating Cloud Deployment for Deep Learning Inference of Real-time Online Services
Yang Li, Zhenhua Han, Quanlu Zhang, Zhenhua Li, Haisheng Tan
摘要
Real-time online services using pre-trained deep neural network (DNN) models, e.g., Siri and Instagram, require low-latency and cost-efficiency for quality-of-service and commercial competitiveness. When deployed in a cloud environment, such services call for an appropriate selection of cloud configurations (i.e., specific types of VM instances), as well as a considerate device placement plan that places the operations of a DNN model to multiple computation devices like GPUs and CPUs. Currently, the deployment mainly relies on service providers' manual efforts, which is not only onerous but also far from satisfactory oftentimes (for a same service, a poor deployment can incur significantly more costs by tens of times). In this paper, we attempt to automate the cloud deployment for real-time online DNN inference with minimum costs under the constraint of acceptably low latency. This attempt is enabled by jointly leveraging the Bayesian Optimization and Deep Reinforcement Learning to adaptively unearth the (nearly) optimal cloud configuration and device placement with limited search time. We implement a prototype system of our solution based on TensorFlow and conduct extensive experiments on top of Microsoft Azure. The results show that our solution essentially outperforms the nontrivial baselines in terms of inference speed and cost-efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- RIBBON: cost-effective and qos-aware deep learning model inference using a diverse pool of cloud computing instancesBaolin Li, Rohan Basu Roy, Tirthak Patel, Vijay Gadepally 等SC 2021 · 被引用 16 次
- ABS: Adaptive Buffer Sizing via Augmented Programmability with Machine LearningJiaxin Tang, Sen Liu, Yang Xu, Zehua Guo 等INFOCOM 2022 · 被引用 9 次
相关 Paper
- Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and SplittingZhixin Zhao, Yitao Hu, Ziqi Gong, Guotao Yang 等INFOCOM 2025 · 被引用 2 次
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo 等UbiComp 2020 · 被引用 77 次
- AoDNN: An Auto-Offloading Approach to Optimize Deep Inference for Fostering Mobile WebYakun Huang, Xiuquan Qiao, Schahram Dustdar, Yan LiINFOCOM 2022 · 被引用 17 次
- Do the Best Cloud Configurations Grow on Trees? An Experimental Evaluation of Black Box Algorithms for Optimizing Cloud Workloads SubMuhammad Bilal, Marco Serafini, Marco Canini, Rodrigo RodriguesVLDB 2020
- Distributed Inference Acceleration with Adaptive DNN Partitioning and OffloadingThaha Mohammed, Carlee Joe-Wong, Rohit Babbar, Mario Di FrancescoINFOCOM 2020 · 被引用 213 次
