Past-Future Scheduler for LLM Serving under SLA Guarantees
Ruihao Gong, Shihao Bai, Siyu Wu, Yunqian Fan, Zaijun Wang, Xiuhong Li, Hailong Yang, Xianglong Liu
摘要
The exploration and application of Large Language Models (LLMs) is thriving. To reduce deployment costs, continuous batching has become an essential feature in current service frameworks. The effectiveness of continuous batching relies on an accurate estimate of the memory requirements of requests. However, due to the diversity in request output lengths, existing frameworks tend to adopt aggressive or conservative schedulers, which often result in significant overestimation or underestimation of memory consumption. Consequently, they suffer from harmful request evictions or prolonged queuing times, failing to achieve satisfactory throughput under strict Service Level Agreement (SLA) guarantees (a.k.a. goodput), across various LLM application scenarios with differing input-output length distributions. To address this issue, we propose a novel Past-Future scheduler that precisely estimates the peak memory resources required by the running batch via considering the historical distribution of request output lengths and calculating memory occupancy at each future time point. It adapts to applications with all types of input-output length distributions, balancing the trade-off between request queuing and harmful evictions, thereby consistently achieving better goodput. Furthermore, to validate the effectiveness of the proposed scheduler, we developed a high-performance LLM serving framework, LightLLM, that implements the Past-Future scheduler. Compared to existing aggressive or conservative schedulers, LightLLM demonstrates superior goodput, achieving up to 2-3× higher goodput than other schedulers under heavy loads. LightLLM is open source to boost the research in such direction (https://github.com/ModelTC/lightllm).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Paper Circle: An Open-source Multi-agent Research Discovery and Analysis FrameworkKomal Kumar, Aman Chadha, Salman Khan, Fahad Shahbaz Khan 等ACL 2026 · 被引用 1 次
- RepetitionCurse: Measuring and Understanding Router Imbalance in Mixture-of-Experts LLMs under DoS StressRuixuan Huang, Qingyue Wang, Hantao Huang, Yudong Gao 等ICML 2026
- Gram2Token: Enabling Run-time GPU-Native Grammar-Constrained Decoding for LLMsHantao Hua, Jiming Su, hao tang, Yiping Yao 等ICML 2026
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQLYou Peng, Youhe Jiang, Wenqi Jiang, Chen Wang 等ICDE 2026
- E^2-SCI: Elastic Edge–Cloud Speculative Decoding via Credit InertiaSenyao Li, Haozhao Wang, Zhaobai Jiang, Zhanbo Jin 等CVPR 2026
它引用的顶会 Paper10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- S3: Increasing GPU Utilization during Generative Inference for Higher ThroughputYunho Jin, Chun-Feng Wu, David Brooks, Gu-Yeon WeiNeurIPS 2023 · 被引用 150 次
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae 等HPDC 2026
- Efficient LLM Scheduling by Learning to RankYichao Fu, Siqi Zhu, Runlong Su, Aurick Qiao 等NeurIPS 2024 · 被引用 129 次
- PreServe: Intelligent Management for LMaaS Systems via Hierarchical PredictionZhihan Jiang, Yujie Huang, Guangba Yu, Junjie Huang 等ICSE 2026 · 被引用 5 次
- MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model ServingTiancheng Zhang, Yulin Chen, Yunfeng Zhao, Shaoyuan Huang 等ICML 2026
