Automating Cloud Deployment for Deep Learning Inference of Real-time Online Services
Yang Li, Zhenhua Han, Quanlu Zhang, Zhenhua Li, Haisheng Tan
Abstract
Real-time online services using pre-trained deep neural network (DNN) models, e.g., Siri and Instagram, require low-latency and cost-efficiency for quality-of-service and commercial competitiveness. When deployed in a cloud environment, such services call for an appropriate selection of cloud configurations (i.e., specific types of VM instances), as well as a considerate device placement plan that places the operations of a DNN model to multiple computation devices like GPUs and CPUs. Currently, the deployment mainly relies on service providers' manual efforts, which is not only onerous but also far from satisfactory oftentimes (for a same service, a poor deployment can incur significantly more costs by tens of times). In this paper, we attempt to automate the cloud deployment for real-time online DNN inference with minimum costs under the constraint of acceptably low latency. This attempt is enabled by jointly leveraging the Bayesian Optimization and Deep Reinforcement Learning to adaptively unearth the (nearly) optimal cloud configuration and device placement with limited search time. We implement a prototype system of our solution based on TensorFlow and conduct extensive experiments on top of Microsoft Azure. The results show that our solution essentially outperforms the nontrivial baselines in terms of inference speed and cost-efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a74c06e0-84c4-4480-adf8-db1523728cf2Cited by top-tier papers2
- RIBBON: cost-effective and qos-aware deep learning model inference using a diverse pool of cloud computing instancesBaolin Li, Rohan Basu Roy, Tirthak Patel, Vijay Gadepally et al.SC 2021 · 16 citations
- ABS: Adaptive Buffer Sizing via Augmented Programmability with Machine LearningJiaxin Tang, Sen Liu, Yang Xu, Zehua Guo et al.INFOCOM 2022 · 9 citations
Related papers
- Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and SplittingZhixin Zhao, Yitao Hu, Ziqi Gong, Guotao Yang et al.INFOCOM 2025 · 2 citations
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo et al.UbiComp 2020 · 77 citations
- AoDNN: An Auto-Offloading Approach to Optimize Deep Inference for Fostering Mobile WebYakun Huang, Xiuquan Qiao, Schahram Dustdar, Yan LiINFOCOM 2022 · 17 citations
- Do the Best Cloud Configurations Grow on Trees? An Experimental Evaluation of Black Box Algorithms for Optimizing Cloud Workloads SubMuhammad Bilal, Marco Serafini, Marco Canini, Rodrigo RodriguesVLDB 2020
- Distributed Inference Acceleration with Adaptive DNN Partitioning and OffloadingThaha Mohammed, Carlee Joe-Wong, Rohit Babbar, Mario Di FrancescoINFOCOM 2020 · 213 citations
