Efficient LLM Scheduling by Learning to Rank
Yichao Fu, Siqi Zhu, Runlong Su, Aurick Qiao, Ion Stoica, Hao Zhang
摘要
In Large Language Model (LLM) inference, the output length of an LLM request is typically regarded as not known a priori. Consequently, most LLM serving systems employ a simple First-come-first-serve (FCFS) scheduling strategy, leading to Head-Of-Line (HOL) blocking and reduced throughput and service quality. In this paper, we reexamine this assumption -- we show that, although predicting the exact generation length of each request is infeasible, it is possible to predict the relative ranks of output lengths in a batch of requests, using learning to rank. The ranking information offers valuable guidance for scheduling requests. Building on this insight, we develop a novel scheduler for LLM inference and serving that can approximate the shortest-job-first (SJF) schedule better than existing approaches. We integrate this scheduler with the state-of-the-art LLM serving system and show significant performance improvement in several important applications: 2.8x lower latency in chatbot serving and 6.5x higher throughput in synthetic data generation. Our code is available at https://github.com/hao-ai-lab/vllm-ltr.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Efficiently Scaling LLM Reasoning Programs with CertaindexYichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu 等NeurIPS 2025 · 被引用 42 次
- Seer: Online Context Learning for Fast Synchronous LLM Reinforcement LearningRuoyu Qin, Weiran He, Weixiao Huang, Yangkun Zhang 等OSDI 2026 · 被引用 32 次
- HyGen: Efficient LLM Serving via Elastic Online-Offline Request Co-locationTing Sun, Penghan Wang, Fan LaiNeurIPS 2025 · 被引用 17 次
- Fast Inference for Augmented Large Language ModelsRana Shahout, Cong Liang, Shiji Xin, Qianru Lao 等NeurIPS 2025 · 被引用 13 次
- Predicting LLM Output Length via Entropy-Guided RepresentationsHuanyi Xie, Yubin Chen, Liangyu Wang, Lijie Hu 等ICLR 2026 · 被引用 12 次
它引用的顶会 Paper9
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- AlpacaFarm: A Simulation Framework for Methods that Learn from Human FeedbackYann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang 等NeurIPS 2023 · 被引用 948 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan 等OSDI 2024 · 被引用 537 次
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li 等ICLR 2024 · 被引用 419 次
相关 Paper
- Scheduling LLM Inference with Uncertainty-Aware Output Length PredictionsHaoyu Zheng, Yongqiang Zhang, Fangcheng Fu, Xiaokai Zhou 等ICML 2026 · 被引用 2 次
- FastServe: Iteration-Level Preemptive Scheduling for Large Language Model InferenceBingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu 等NSDI 2026 · 被引用 12 次
- Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference PipelineZangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo 等NeurIPS 2023 · 被引用 159 次
- Don't stop me Now: Embedding based Scheduling for LLMSRana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang 等ICLR 2025
- S3: Increasing GPU Utilization during Generative Inference for Higher ThroughputYunho Jin, Chun-Feng Wu, David Brooks, Gu-Yeon WeiNeurIPS 2023 · 被引用 150 次
