InfScaler: Enabling Efficient ML Inference Serving on Multi-Accelerator Edge Devices via Asymmetric Auto-Scaling
Borui Li, Tiange Xia, Shuai Wang, Shuai Wang
摘要
Nowadays, there is a growing trend to deploy machine learning (ML) models on edge devices. To cope with the increasing resource requirements of current ML models, multi-accelerator edge devices that integrate CPU, GPU, NPU, or TPU in a single SoC gain popularity. However, we observe that existing ML inference serving frameworks are poor in utilizing the unique hardware architecture of these edge devices. In this paper, we present InFSCALER, an efficient ML inference serving framework tailored for multi-accelerator edge devices. InfSCALER discovers the architectural bottleneck of ML models and designs a bottleneck-aware asymmetric auto-scaling technique to facilitate efficient resource allocation for ML models on the edge. Furthermore, InfScaler capitalizes on the hardware’s unified memory feature inherent to edge devices, ensuring efficient data sharing between the asymmetrically scaled model partitions. Our experimental results show that InfScaler achieves up to throughput improvement and resource reduction while satisfying the latency requirements compared with the state-of-the-art inference serving approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 被引用 184 次
- INFless: a native serverless system for low-latency, high-throughput inferenceYanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang 等ASPLOS 2022 · 被引用 145 次
- VIPS: real-time perception fusion for infrastructure-assisted autonomous drivingShuyao Shi, Jiahe Cui, Zhehao Jiang, Zhenyu Yan 等MobiCom 2022 · 被引用 126 次
- SPRIGHT: extracting the server from serverless computing! high-performance eBPF-based event-driven, shared-memory processingShixiong Qi, Leslie Monis, Ziteng Zeng, Ian-Chin Wang 等SIGCOMM 2022 · 被引用 85 次
相关 Paper
- Minimizing Latency for Multi-DNN Inference on Resource-Limited CPU-Only Edge DevicesTao Wang, Tuo Shi, Xiulong Liu, Jianping Wang 等INFOCOM 2024 · 被引用 9 次
- AsyMo: scalable and efficient deep-learning inference on asymmetric mobile CPUsManni Wang, Shaohua Ding, Ting Cao, Yunxin Liu 等MobiCom 2021 · 被引用 69 次
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao 等HPCA 2026
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park 等USENIX ATC 2022 · 被引用 200 次
- EdgeNN: Efficient Neural Network Inference for CPU-GPU Integrated Edge DevicesChenyang Zhang, Feng Zhang, Kuangyu Chen, Mingjun Chen 等ICDE 2023 · 被引用 16 次
