Lune

EuroSys2025Top-tier venue

Flex: Fast, Accurate DNN Inference on Low-Cost Edges Using Heterogeneous Accelerator Execution

Tanmoy Sen, Haiying Shen, Anand Padmanabha Iyer

2025Year
2Citations

Abstract

Significant b reakthroughs in machine learning (ML) and the advantages of on-device processing have led to edge devices increasingly incorporating accelerators like GPUs, NPUs, and DSPs. However, these accelerators consume energy, prompting users to limit their floating-point precision. Many edge device users are in regions where including high-fidelity accelerators is too costly, leading to low-cost devices with low precision, sacrificing accuracy. Previous work predetermined layer assignments between the CPU and accelerator offline for high accuracy and low latency without considering the input, but we observe that input affects optimal layer assignment. To address this, we present Flex, a system for Fast, Accurate DNN Inference on Low-Cost Edges using Heterogeneous Accelerator eXecution. Leveraging common observations from models on various edge devices, Flex uses a lightweight heuristic and reinforcement learning (RL) to dynamically assign layers across the CPU and accelerator. Experiments show Flex improves average inference time by up to 39%, accuracy by up to 22%, and energy consumption by up to 61% compared to state-of-the-art methods, and is only 4.2% less optimal than the best achievable results.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 866cbd7b-ecf2-486f-9f1b-4182b66f2e64

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines