Lazy but Efficient: Layer-Wise Task Scheduling with Lazy Pulling for Fast Serverless Inference
Zhexiong Li, Hongmin Geng, Yuepeng Li, Lin Gu, Deze Zeng
Abstract
Serverless inference is becoming increasingly popular in AI services thanks to its scalability and flexibility. However, high latency caused by pulling large container images and AI models remains a significant bottleneck. Lazy pulling, which pulls only the specific image layers and model layers when needed, helps reduce this latency by allowing the pulling process to overlap with inference execution in a pipeline manner. Nevertheless, we discover that lazy pulling exhibits interwoven dependencies between container image layers and AI model layers. This not only results in unnecessary idle periods where the system waits for the required layers to be pulled, but also complicates layer scheduling. To this end, in this paper, we investigate a Layer-wise Serverless Inference Scheduling (LSIS) problem, and formulate it as an Integer Linear Program (ILP) form. We further propose a LayerChain algorithm that utilizes a Variable-Length Time Slot (VL-Slot) and provides a theoretical analysis of its performance upper bound. Trace-driven evaluations demonstrate that our approach significantly reduces inference completion time compared with state-of-the-art methods by 38.49%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bc90a6c0-4c75-4b10-b79c-45f599d77659Related papers
- FaSei: Fast Serverless Edge Inference with Synergistic Lazy Loading and Layer-wise CachingZhaowu Huang, Fang Dong, Xiaolin Guo, Daheng YinINFOCOM 2025 · 4 citations
- Optimus: Warming Serverless ML Inference via Inter-Function Model TransformationZicong Hong, Jian Lin, Song Guo, Sifu Luo et al.EuroSys 2024 · 29 citations
- Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient InferenceMinchen Yu, Ao Wang, Dong Chen, Haoxuan Yu et al.USENIX ATC 2025
- Towards Resource-Efficient Serverless LLM Inference with SLINFERChuhao Xu, Zijun Li, Quan Chen, Han Zhao et al.HPCA 2026
- ServerlessLLM: Low-Latency Serverless Inference for Large Language ModelsYao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete et al.OSDI 2024 · 125 citations
