Janus: Multi-LLM Serving at Production Scale
Tianbao Zhou, Yi Wang, Yu Zhou, Zirui Liu, Zhiming Wang, Yebo Peng, Yongfu Wang, Yi Zhang, Jinrun Yin, Kemeng Tian, Fangcheng Fu, Tongxuan Liu
Abstract
Production LLM serving multiplexes hundreds of heterogeneous models on shared clusters, exposing three challenges that existing systems fail to address simultaneously: unpredictable bursts, power-law application popularity, and heterogeneous yet complementary resource demands. We present Janus, a Service-Engine co-designed multi-model serving system built around dual-timescale scheduling. At the Service layer, a Model Scheduler driven by a Performance Oracle combines vector bin-packing every 30 seconds for steady-state colocation with elastic scaling every 0.5 seconds. A new LST-IMH Request Scheduler maximizes SLO attainment with a proven 2-approximation guarantee. At the Engine layer, xTensor virtualizes HBM, while a three-state model lifecycle and device-to-device (D2D) fork enable elastic multi-model colocation and sub-second burst scale-out. We deploy Janus on a 768-device production cluster serving 62 applications and 45.6 M requests/day, and evaluate it with a 603 K-request open-loop replay sampled from that production day. Janus maintains 0.97-1.0 SLO attainment, versus 0.80-0.92 for the strongest baseline, while using 13,440 device-hours/day, 27% less than static Serverless-LLM. Janus is available as open source at https://github.com/Janus2026/Janus.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 64f4cc61-97c8-4c0a-92de-b2e0487a8ae9Related papers
- MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM ServingJiangfei Duan, Runyu Lu, Haojie Duanmu, Xiuhong Li et al.ICML 2024 · 51 citations
- Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM WorkloadsChaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi et al.NSDI 2026 · 5 citations
- Laser: Unlocking Layer-Level Scheduling for Efficient Multi-SLO LLM ServingJianxiong Liao, Quanxing Dong, Yunkai Liang, Zhi Zhou et al.PPoPP 2026 · 1 citation
- Llumnix: Dynamic Scheduling for Large Language Model ServingBiao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao et al.OSDI 2024 · 189 citations
- WarmServe: Enabling One-for-Many GPU Prewarming for Multi-LLM ServingChiheng Lou, Sheng Qi, Rui Kang, Yong Zhang et al.ICML 2026 · 3 citations
