Lune

INFOCOM2026顶会

Lazy but Efficient: Layer-Wise Task Scheduling with Lazy Pulling for Fast Serverless Inference

Zhexiong Li, Hongmin Geng, Yuepeng Li, Lin Gu, Deze Zeng

2026年份
1被引次数

摘要

Serverless inference is becoming increasingly popular in AI services thanks to its scalability and flexibility. However, high latency caused by pulling large container images and AI models remains a significant bottleneck. Lazy pulling, which pulls only the specific image layers and model layers when needed, helps reduce this latency by allowing the pulling process to overlap with inference execution in a pipeline manner. Nevertheless, we discover that lazy pulling exhibits interwoven dependencies between container image layers and AI model layers. This not only results in unnecessary idle periods where the system waits for the required layers to be pulled, but also complicates layer scheduling. To this end, in this paper, we investigate a Layer-wise Serverless Inference Scheduling (LSIS) problem, and formulate it as an Integer Linear Program (ILP) form. We further propose a LayerChain algorithm that utilizes a Variable-Length Time Slot (VL-Slot) and provides a theoretical analysis of its performance upper bound. Trace-driven evaluations demonstrate that our approach significantly reduces inference completion time compared with state-of-the-art methods by 38.49%.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖