Lune

INFOCOM2026Top-tier venue

Rocket: Warming Serverless Inference via Hierarchical ML Artifact Pre-loading and Sharing

Xiaofei Yue, Song Yang, Fan Li, Youqi Li, Yu Wang

2026Year
2Citations

Abstract

Serverless computing is a promising method to serve Machine Learning (ML) inference via on-demand functions. Due to the time- and memory-consuming ML library and model (i.e., ML artifact) loading, serverless inference endures notable startup overhead and memory waste issues. In this paper, we advocate for hierarchical ML artifact pre-loading and sharing to balance loading and memory efficiency. Building on this, we propose Rocket, a serverless ML inference system that accelerates function startup while reducing memory waste. Rocket dynamically pre-loads partial, shared ML artifacts, each implying a hierarchy of trade-offs between the loading latency and memory usage. Specifically, with a dual-timescale invocation prediction, Rocket first estimates the pre-loading timing for each function, and then schedules them via a sharing-aware agglomerative clustering to improve ML artifact sharing efficiency. In particular, Rocket learns to make the online hierarchical pre-loading decision for function containers based on a lightweight contextual bandit algorithm. Finally, we implement Rocket and evaluate it with realistic workloads. Experimental results display that Rocket outperforms existing solutions by up to 38.7% on startup latency and up to 43.8% on memory saving.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get 5ea693f7-0c66-4a17-a77f-917e053a51f7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines