Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures
Yongji Wu, Matthew Lentz, Danyang Zhuo, Yao Lu
Abstract
With the advent of ubiquitous deployment of smart devices and the Internet of Things, data sources for machine learning inference have increasingly moved to the edge of the network. Existing machine learning inference platforms typically assume a homogeneous infrastructure and do not take into account the more complex and tiered computing infrastructure that includes edge devices, local hubs, edge datacenters, and cloud datacenters. On the other hand, recent AutoML efforts have provided viable solutions for model compression, pruning and quantization for heterogeneous environments; for a machine learning model, now we may easily find or even generate a series of model variants with different tradeoffs between accuracy and efficiency. We design and implement JellyBean, a system for serving and optimizing machine learning inference workflows on heterogeneous infrastructures. Given service-level objectives (e.g., throughput, accuracy), JellyBean picks the most cost-efficient models that meet the accuracy target and decides how to deploy them across different tiers of infrastructures. Evaluations show that JellyBean reduces the total serving cost of visual question answering by up to 58% and vehicle tracking from the NVIDIA AI City Challenge by up to 36%, compared with state-of-the-art model selection and worker assignment solutions. JellyBean also outperforms prior ML serving systems (e.g., Spark on the cloud) up to 5x in serving costs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Optimizing Video Analytics with Declarative Model RelationshipsFrancisco Romero, Johann Hauswald, Aditi Partap, Daniel Kang et al.VLDB 2023 · 37 citations
- DAHA: Accelerating GNN Training with Data and Hardware Aware Execution PlanningZhiyuan Li, Xun Jian, Yue Wang, Yingxia Shao et al.VLDB 2024 · 18 citations
- Vulcan: Automatic Query Planning for Live ML AnalyticsYiwen Zhang, Xumiao Zhang, Ganesh Ananthanarayanan, Anand P. Iyer et al.NSDI 2024 · 17 citations
- Compass: SLO-aware Query Planner for Compound AI Serving at ScaleBanruo Liu, Wei-Yu Lin, Minghao Fang, Yihan Jiang et al.VLDB 2026 · 5 citations
- OTAS: An Elastic Transformer Serving System via Token AdaptationJinyu Chen, Wenchao Xu, Zicong Hong, Song Guo et al.INFOCOM 2024 · 4 citations
Builds on7
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
- Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video AnalyticsYuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang et al.SIGCOMM 2020 · 264 citations
- Joint Configuration Adaptation and Bandwidth Allocation for Edge-based Real-time Video AnalyticsCan Wang, Sheng Zhang, Yu Chen, Zhuzhong Qian et al.INFOCOM 2020 · 223 citations
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia et al.MobiCom 2021 · 171 citations
Related papers
- RIBBON: cost-effective and qos-aware deep learning model inference using a diverse pool of cloud computing instancesBaolin Li, Rohan Basu Roy, Tirthak Patel, Vijay Gadepally et al.SC 2021 · 16 citations
- TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge DevicesBowen Zhang, Junyang Zhang, Jiahui Hou, Yixin WangINFOCOM 2025 · 10 citations
- SmartLite: A DBMS-based Serving System for DNN Inference in Resource-constrained EnvironmentsQiuru Lin, Sai Wu, Junbo Zhao, Jian Dai et al.VLDB 2024 · 17 citations
- Proteus: A High-Throughput Inference-Serving System with Accuracy ScalingSohaib Ahmad, Hui Guan, Brian D. Friedman, Thomas Williams et al.ASPLOS 2024 · 31 citations
- Joint Model and Data Adaptation for Cloud Inference ServingJingyan Jiang, Ziyue Luo, Chenghao Hu, Zhaoliang He et al.RTSS 2021 · 19 citations
