FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise Caching
Zhaowu Huang, Fang Dong, Xiaolin Guo, Daheng Yin
摘要
Serverless edge computing (SEC) provides low-latency, resource-efficient deep learning (DL) services but faces a significant cold start time due to the loading of large DL models. Existing methods for cold starts include full and partial container caching. The former may be inefficient because large DL models cannot be cached in the resource-limited SEC; The latter only caches common packages for sharing, while user-specified DL models still need to be loaded before execution. We identify model lazy loading to mitigate cold start, which begins inference with a shallow model while lazily loading deeper layers in a pipeline manner. However, naively using lazy loading can result in significant inference bubbles because model loading time is typically much longer than inference time. To address this, we propose FaSei, a fast serverless edge inference method with synergistic model lazy loading and layer-wise caching, to reduce application completion time (ACT). Considering the impact of heterogeneous model-layer behaviors and SEC resources on ACT, we jointly optimize lazy loading, layer-wise caching, and function placement. We formulate it as an integer nonlinear programming problem and then design an approximation algorithm with a theoretical performance guarantee. Extensive experiments demonstrate that FaSei achieves up tospeedup in reducing ACT.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Online Container Caching with Late-Warm for IoT Data ProcessingGuopeng Li, Haisheng Tan, Xuan Zhang, Chi Zhang 等ICDE 2024 · 被引用 4 次
- Lazy but Efficient: Layer-Wise Task Scheduling with Lazy Pulling for Fast Serverless InferenceZhexiong Li, Hongmin Geng, Yuepeng Li, Lin Gu 等INFOCOM 2026 · 被引用 1 次
- RainbowCake: Mitigating Cold-starts in Serverless with Layer-wise Container Caching and SharingHanfei Yu, Rohan Basu Roy, Christian Fontenot, Devesh Tiwari 等ASPLOS 2024 · 被引用 69 次
- On Efficient Zygote Container Planning toward Fast Function Startup in Serverless Edge CloudYuepeng Li, Deze Zeng, Lin Gu, Mingwei Ou 等INFOCOM 2023 · 被引用 23 次
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo 等EuroSys 2024 · 被引用 29 次
