Lune

DAC2025Top-tier venue

InfScaler: Enabling Efficient ML Inference Serving on Multi-Accelerator Edge Devices via Asymmetric Auto-Scaling

Borui Li, Tiange Xia, Shuai Wang, Shuai Wang

2025Year
2Citations

Abstract

Nowadays, there is a growing trend to deploy machine learning (ML) models on edge devices. To cope with the increasing resource requirements of current ML models, multi-accelerator edge devices that integrate CPU, GPU, NPU, or TPU in a single SoC gain popularity. However, we observe that existing ML inference serving frameworks are poor in utilizing the unique hardware architecture of these edge devices. In this paper, we present InFSCALER, an efficient ML inference serving framework tailored for multi-accelerator edge devices. InfSCALER discovers the architectural bottleneck of ML models and designs a bottleneck-aware asymmetric auto-scaling technique to facilitate efficient resource allocation for ML models on the edge. Furthermore, InfScaler capitalizes on the hardware’s unified memory feature inherent to edge devices, ensuring efficient data sharing between the asymmetrically scaled model partitions. Our experimental results show that InfScaler achieves up to 126.59%\mathbf{1 2 6. 5 9 \%} throughput improvement and 27.32%\mathbf{2 7. 3 2 \%} resource reduction while satisfying the latency requirements compared with the state-of-the-art inference serving approaches.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext c778fec1-0894-481a-8f9f-ccdda9dd3913

Builds on19

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines