Don't stop me Now: Embedding based Scheduling for LLMS
Rana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang, Minlan Yu, Michael Mitzenmacher
摘要
Efficient scheduling is crucial for interactive Large Language Model (LLM) applications, where low request completion time directly impacts user engagement. Size-based scheduling algorithms like Shortest Remaining Process Time (SRPT) aim to reduce average request completion time by leveraging known or estimated request sizes and allowing preemption by incoming jobs with shorter service times. However, two main challenges arise when applying size-based scheduling to LLM systems. First, accurately predicting output lengths from prompts is challenging and often resource-intensive, making it impractical for many systems. As a result, the state-of-the-art LLM systems default to first-come, first-served scheduling, which can lead to head-of-line blocking and reduced system efficiency. Second, preemption introduces extra memory overhead to LLM systems as they must maintain intermediate states for unfinished (preempted) requests. In this paper, we propose TRAIL, a method to obtain output predictions from the target LLM itself. After generating each output token, we recycle the embedding of its internal structure as input for a lightweight classifier that predicts the remaining length for each running request. Using these predictions, we propose a prediction-based SRPT variant with limited preemption designed to account for memory overhead in LLM systems. This variant allows preemption early in request execution when memory consumption is low but restricts preemption as requests approach completion to optimize resource utilization. On the theoretical side, we derive a closed-form formula for this SRPT variant in an M/G/1 queue model, which demonstrates its potential value. In our system, we implement this preemption policy alongside our embedding-based prediction method. Our refined predictions from layer embeddings achieve 2.66x lower mean absolute error compared to BERT predictions from sequence prompts. TRAIL achieves 1.66x to 2.01x lower mean latency on the Alpaca dataset and 1.76x to 24.07x lower mean time to the first token compared to the state-of-the-art serving system.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Fast Inference for Augmented Large Language ModelsRana Shahout, Cong Liang, Shiji Xin, Qianru Lao 等NeurIPS 2025 · 被引用 13 次
- Predicting LLM Output Length via Entropy-Guided RepresentationsHuanyi Xie, Yubin Chen, Liangyu Wang, Lijie Hu 等ICLR 2026 · 被引用 12 次
- Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM WorkloadsChaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi 等NSDI 2026 · 被引用 5 次
- AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference ServingYing Wang, Zhen Jin, Zhenqian Chen, Jiexiong Xu 等ICML 2026 · 被引用 4 次
- Scheduling LLM Inference with Uncertainty-Aware Output Length PredictionsHaoyu Zheng, Yongqiang Zhang, Fangcheng Fu, Xiaokai Zhou 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper5
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference PipelineZangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo 等NeurIPS 2023 · 被引用 159 次
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas 等HPCA 2025 · 被引用 106 次
- SkipPredict: When to Invest in Predictions for SchedulingRana Shahout, Michael MitzenmacherNeurIPS 2024 · 被引用 6 次
相关 Paper
- Efficient LLM Scheduling by Learning to RankYichao Fu, Siqi Zhu, Runlong Su, Aurick Qiao 等NeurIPS 2024 · 被引用 129 次
- Beyond Prediction: Tail-Aware Scheduling for LLM InferenceYueying Li, Yuanfan Chen, Jiayang Chen, Esha Choukse 等ICML 2026
- FastServe: Iteration-Level Preemptive Scheduling for Large Language Model InferenceBingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu 等NSDI 2026 · 被引用 12 次
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae 等HPDC 2026
- S3: Increasing GPU Utilization during Generative Inference for Higher ThroughputYunho Jin, Chun-Feng Wu, David Brooks, Gu-Yeon WeiNeurIPS 2023 · 被引用 150 次
