AdaGen: Workload-Adaptive Cluster Scheduler for Latency-Optimal LLM Inference Serving
Sudipta Saha Shubha, Ayush Goel, Diman Zad Tootaghaj, Khaled Diab, Hardik Soni, K. K. Ramakrishnan, Puneet Sharma, Haiying Shen
Abstract
The inference workloads of Large Language Models (LLMs) pose significant latency and cost challenges due to increasing model sizes and demand for real-time responses. Existing cluster schedulers for multi-instance LLM serving primarily focus on load balancing to optimize memory usage, which is insufficient for workloads with diverse request characteristics. In such cases, the compute layout — the arrangement of tokens across iterations within each instance—plays a crucial role in determining latency. We propose AdaGen, a workload-adaptive cluster scheduler that minimizes latency and thus maximizes SLO attainment by optimizing compute layouts across instances. AdaGen employs a multi-step scheduling strategy: it first classifies requests based on prefill and decode lengths, then balances load, and finally performs selective distributed execution across instances. Each step incrementally refines the scheduling based on the compute layouts derived from the decision of the previous step. To avoid the overhead of actual execution to generate the layouts, AdaGen introduces a novel simulation-based estimator. Extensive experiments using production workloads show that AdaGen achieves up to 3.6× higher SLO attainment and 2× better cost-efficiency compared to the existing systems, while ensuring scalability.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 06649a2d-ef22-4b95-a8d6-0a677847807bRelated papers
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQLYou Peng, Youhe Jiang, Wenqi Jiang, Chen Wang et al.ICDE 2026
- Efficient Multi-round LLM Inference over Disaggregated ServingWenhao He, Youhe Jiang, Penghao Zhao, Quanqing Xu et al.ICML 2026 · 7 citations
- AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference ServingYing Wang, Zhen Jin, Zhenqian Chen, Jiexiong Xu et al.ICML 2026 · 4 citations
- PreServe: Intelligent Management for LMaaS Systems via Hierarchical PredictionZhihan Jiang, Yujie Huang, Guangba Yu, Junjie Huang et al.ICSE 2026 · 5 citations
- HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-locationTing Sun, Penghan Wang, Fan LaiNeurIPS 2025 · 17 citations
