Lune

INFOCOM2025Top-tier venue

FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise Caching

Zhaowu Huang, Fang Dong, Xiaolin Guo, Daheng Yin

2025Year
4Citations

Abstract

Serverless edge computing (SEC) provides low-latency, resource-efficient deep learning (DL) services but faces a significant cold start time due to the loading of large DL models. Existing methods for cold starts include full and partial container caching. The former may be inefficient because large DL models cannot be cached in the resource-limited SEC; The latter only caches common packages for sharing, while user-specified DL models still need to be loaded before execution. We identify model lazy loading to mitigate cold start, which begins inference with a shallow model while lazily loading deeper layers in a pipeline manner. However, naively using lazy loading can result in significant inference bubbles because model loading time is typically much longer than inference time. To address this, we propose FaSei, a fast serverless edge inference method with synergistic model lazy loading and layer-wise caching, to reduce application completion time (ACT). Considering the impact of heterogeneous model-layer behaviors and SEC resources on ACT, we jointly optimize lazy loading, layer-wise caching, and function placement. We formulate it as an integer nonlinear programming problem and then design an approximation algorithm with a theoretical performance guarantee. Extensive experiments demonstrate that FaSei achieves up to6.7×6.7\timesspeedup in reducing ACT.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get ef8a80ca-7a45-4c0d-9e7a-4203e8933052

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines