InfScaler: Enabling Efficient ML Inference Serving on Multi-Accelerator Edge Devices via Asymmetric Auto-Scaling
Borui Li, Tiange Xia, Shuai Wang, Shuai Wang
Abstract
Nowadays, there is a growing trend to deploy machine learning (ML) models on edge devices. To cope with the increasing resource requirements of current ML models, multi-accelerator edge devices that integrate CPU, GPU, NPU, or TPU in a single SoC gain popularity. However, we observe that existing ML inference serving frameworks are poor in utilizing the unique hardware architecture of these edge devices. In this paper, we present InFSCALER, an efficient ML inference serving framework tailored for multi-accelerator edge devices. InfSCALER discovers the architectural bottleneck of ML models and designs a bottleneck-aware asymmetric auto-scaling technique to facilitate efficient resource allocation for ML models on the edge. Furthermore, InfScaler capitalizes on the hardware’s unified memory feature inherent to edge devices, ensuring efficient data sharing between the asymmetrically scaled model partitions. Our experimental results show that InfScaler achieves up to throughput improvement and resource reduction while satisfying the latency requirements compared with the state-of-the-art inference serving approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c778fec1-0894-481a-8f9f-ccdda9dd3913Builds on19
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 184 citations
- INFless: a native serverless system for low-latency, high-throughput inferenceYanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang et al.ASPLOS 2022 · 145 citations
- VIPS: real-time perception fusion for infrastructure-assisted autonomous drivingShuyao Shi, Jiahe Cui, Zhehao Jiang, Zhenyu Yan et al.MobiCom 2022 · 126 citations
- SPRIGHT: extracting the server from serverless computing! high-performance eBPF-based event-driven, shared-memory processingShixiong Qi, Leslie Monis, Ziteng Zeng, Ian-Chin Wang et al.SIGCOMM 2022 · 85 citations
Related papers
- Minimizing Latency for Multi-DNN Inference on Resource-Limited CPU-Only Edge DevicesTao Wang, Tuo Shi, Xiulong Liu, Jianping Wang et al.INFOCOM 2024 · 9 citations
- AsyMo: scalable and efficient deep-learning inference on asymmetric mobile CPUsManni Wang, Shaohua Ding, Ting Cao, Yunxin Liu et al.MobiCom 2021 · 69 citations
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao et al.HPCA 2026
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- EdgeNN: Efficient Neural Network Inference for CPU-GPU Integrated Edge DevicesChenyang Zhang, Feng Zhang, Kuangyu Chen, Mingjun Chen et al.ICDE 2023 · 16 citations
