Simple Is Better: Multiplication May Be All You Need for LLM Request Scheduling
Dingyan Zhang, Jinbo Han, Kaixi Zhang, Xingda Wei, Sijie Shen, Chenguang Fang, Wenyuan Yu, Jingren Zhou, Rong Chen
Abstract
High-quality LLM request scheduling requires meeting two key objectives: ensuring the routed instance has KV$ to accelerate request execution, and ensuring that the workload is balanced across instances. Achieving both objectives is challenging because pursuing one may compromise the other. Current approaches use various combinators (e.g., linear combinations) to compute a scheduling score that combines indicators for the two objectives. These approaches are complex: they either require significant workload-specific hyperparameter tuning or model-hardware-aware simulator development, yet could still lead to suboptimal performance.
In this paper, we show that using a simple multiplication of two carefully chosen indicators-one KV$-aware (new prefill tokens if routed to an instance) and one load-balancing-aware (current batch size of the instance)-as the scheduling score (LMETRIC * ) can achieve both objectives simultaneously without any hyperparameter tuning. The key idea is that the simply multiplied score considers both objectives in a manner similar to a linear combination, but the original hyperparameters cancel out during comparison, so no tuning is needed to find the best parameters. The two indicators are chosen based on our analysis of LLM characteristics. Our extensive experiments show that this simple approach can reduce TTFT by 92% and 39%, and TPOT by 24% and 51%, compared to vLLM-v1 and an in-production scheduler on real-world workloads covering chatbots and coding agents. We also derive the mathematical conditions under which multiplication may fail, and find that such conditions are extremely rare in practice and can be detected (and mitigated) beforehand.
LMETRIC has been deployed in production and canary release confirms its effectiveness. * Our system name LMETRIC stands for Large Model metric, paying homage to Lyapunov's stability theory and Markov's queueing theory. † Most work done when intern at Alibaba Group. ‡ Jinbo Han and Kaixi Zhang contributed equally and significantly to this work.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f750336d-715d-42d9-933a-445aa54934bfBuilds on13
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttentionBin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang et al.USENIX ATC 2024 · 273 citations
Related papers
- DualMap: Enabling Both Cache Affinity and Load Balancing for Distributed LLM ServingYing Yuan, Pengfei Zuo, Bo Wang, Zhangyu Chen et al.ICLR 2026 · 10 citations
- Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM WorkloadsChaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi et al.NSDI 2026 · 5 citations
- LLM Query Scheduling with Prefix Reuse and Latency ConstraintsGregory Dexter, Shao Tang, Ata Fatahi Baarzi, Qingquan Song et al.NeurIPS 2025 · 10 citations
- LLM-as-Scheduler: Agentic Workflow Dynamic SchedulingDawei Xiang, Kexin Chu, Wenyan Xu, Wenhui Zhang et al.ACL 2026
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning et al.NSDI 2026 · 29 citations
