Automated End-to-End Model Serving with Cooperative Compilation and Scheduling
Yikang Zhang, Junlong Chen, Wei Wang, Jia Liu, Nan Hu, Haipeng Dai
Abstract
Model serving systems are critical for deep learning inference, managing GPU infrastructure to deliver end-to-end services. Current frameworks typically treat operators as basic compilation and scheduling units, which often fail to maximize GPU utilization due to hardware-unfriendly kernels and coarse-grained scheduling. To address these limitations, we propose a cooperative compilation and scheduling scheme that statically generates multiple kernel variants and dynamically schedules them based on runtime context. We present Infera, a high-performance model serving system that implements this approach. Experimental results demonstrate that Infera improves inference throughput by at least 1.6× compared to state-of-the-art baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get ef00d2bf-9641-4be9-9c1a-abc084d62f10Related papers
- Paella: Low-latency Model Serving with Software-defined GPU SchedulingKelvin K. W. Ng, Henri Maxime Demoulin, Vincent LiuSOSP 2023 · 30 citations
- USHER: Holistic Interference Avoidance for Resource Optimized ML InferenceSudipta Saha Shubha, Haiying Shen, Anand P. IyerOSDI 2024 · 35 citations
- Compile-Time QoS Scheme for Deep Learning InferencesSungin Hong, Hyunjun Kim, Hwansoo HanSC 2025 · 1 citation
- Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM WorkloadsChaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi et al.NSDI 2026 · 5 citations
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
