Lune

EuroSys2026Top-tier venue

Automated End-to-End Model Serving with Cooperative Compilation and Scheduling

Yikang Zhang, Junlong Chen, Wei Wang, Jia Liu, Nan Hu, Haipeng Dai

2026Year

Abstract

Model serving systems are critical for deep learning inference, managing GPU infrastructure to deliver end-to-end services. Current frameworks typically treat operators as basic compilation and scheduling units, which often fail to maximize GPU utilization due to hardware-unfriendly kernels and coarse-grained scheduling. To address these limitations, we propose a cooperative compilation and scheduling scheme that statically generates multiple kernel variants and dynamically schedules them based on runtime context. We present Infera, a high-performance model serving system that implements this approach. Experimental results demonstrate that Infera improves inference throughput by at least 1.6× compared to state-of-the-art baselines.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get ef00d2bf-9641-4be9-9c1a-abc084d62f10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines