Serving and Optimizing Machine Learning Workflows on Heterogeneous Infrastructures
Yongji Wu, Matthew Lentz, Danyang Zhuo, Yao Lu
摘要
With the advent of ubiquitous deployment of smart devices and the Internet of Things, data sources for machine learning inference have increasingly moved to the edge of the network. Existing machine learning inference platforms typically assume a homogeneous infrastructure and do not take into account the more complex and tiered computing infrastructure that includes edge devices, local hubs, edge datacenters, and cloud datacenters. On the other hand, recent AutoML efforts have provided viable solutions for model compression, pruning and quantization for heterogeneous environments; for a machine learning model, now we may easily find or even generate a series of model variants with different tradeoffs between accuracy and efficiency. We design and implement JellyBean, a system for serving and optimizing machine learning inference workflows on heterogeneous infrastructures. Given service-level objectives (e.g., throughput, accuracy), JellyBean picks the most cost-efficient models that meet the accuracy target and decides how to deploy them across different tiers of infrastructures. Evaluations show that JellyBean reduces the total serving cost of visual question answering by up to 58% and vehicle tracking from the NVIDIA AI City Challenge by up to 36%, compared with state-of-the-art model selection and worker assignment solutions. JellyBean also outperforms prior ML serving systems (e.g., Spark on the cloud) up to 5x in serving costs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Optimizing Video Analytics with Declarative Model RelationshipsFrancisco Romero, Johann Hauswald, Aditi Partap, Daniel Kang 等VLDB 2023 · 被引用 37 次
- DAHA: Accelerating GNN Training with Data and Hardware Aware Execution PlanningZhiyuan Li, Xun Jian, Yue Wang, Yingxia Shao 等VLDB 2024 · 被引用 18 次
- Vulcan: Automatic Query Planning for Live ML AnalyticsYiwen Zhang, Xumiao Zhang, Ganesh Ananthanarayanan, Anand P. Iyer 等NSDI 2024 · 被引用 17 次
- Compass: SLO-aware Query Planner for Compound AI Serving at ScaleBanruo Liu, Wei-Yu Lin, Minghao Fang, Yihan Jiang 等VLDB 2026 · 被引用 5 次
- OTAS: An Elastic Transformer Serving System via Token AdaptationJinyu Chen, Wenchao Xu, Zicong Hong, Song Guo 等INFOCOM 2024 · 被引用 4 次
它引用的顶会 Paper7
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy 等ICLR 2020 · 被引用 1,037 次
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 被引用 656 次
- Reducto: On-Camera Filtering for Resource-Efficient Real-Time Video AnalyticsYuanqi Li, Arthi Padmanabhan, Pengzhan Zhao, Yufei Wang 等SIGCOMM 2020 · 被引用 264 次
- Joint Configuration Adaptation and Bandwidth Allocation for Edge-based Real-time Video AnalyticsCan Wang, Sheng Zhang, Yu Chen, Zhuzhong Qian 等INFOCOM 2020 · 被引用 223 次
- Elf: accelerate high-resolution mobile deep vision with content-aware parallel offloadingWuyang Zhang, Zhezhi He, Luyang Liu, Zhenhua Jia 等MobiCom 2021 · 被引用 171 次
相关 Paper
- RIBBON: cost-effective and qos-aware deep learning model inference using a diverse pool of cloud computing instancesBaolin Li, Rohan Basu Roy, Tirthak Patel, Vijay Gadepally 等SC 2021 · 被引用 16 次
- TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge DevicesBowen Zhang, Junyang Zhang, Jiahui Hou, Yixin WangINFOCOM 2025 · 被引用 10 次
- SmartLite: A DBMS-based Serving System for DNN Inference in Resource-constrained EnvironmentsQiuru Lin, Sai Wu, Junbo Zhao, Jian Dai 等VLDB 2024 · 被引用 17 次
- Proteus: A High-Throughput Inference-Serving System with Accuracy ScalingSohaib Ahmad, Hui Guan, Brian D. Friedman, Thomas Williams 等ASPLOS 2024 · 被引用 31 次
- Joint Model and Data Adaptation for Cloud Inference ServingJingyan Jiang, Ziyue Luo, Chenghao Hu, Zhaoliang He 等RTSS 2021 · 被引用 19 次
