Don't stop me Now: Embedding based Scheduling for LLMS
Rana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang, Minlan Yu, Michael Mitzenmacher
Abstract
Efficient scheduling is crucial for interactive Large Language Model (LLM) applications, where low request completion time directly impacts user engagement. Size-based scheduling algorithms like Shortest Remaining Process Time (SRPT) aim to reduce average request completion time by leveraging known or estimated request sizes and allowing preemption by incoming jobs with shorter service times. However, two main challenges arise when applying size-based scheduling to LLM systems. First, accurately predicting output lengths from prompts is challenging and often resource-intensive, making it impractical for many systems. As a result, the state-of-the-art LLM systems default to first-come, first-served scheduling, which can lead to head-of-line blocking and reduced system efficiency. Second, preemption introduces extra memory overhead to LLM systems as they must maintain intermediate states for unfinished (preempted) requests. In this paper, we propose TRAIL, a method to obtain output predictions from the target LLM itself. After generating each output token, we recycle the embedding of its internal structure as input for a lightweight classifier that predicts the remaining length for each running request. Using these predictions, we propose a prediction-based SRPT variant with limited preemption designed to account for memory overhead in LLM systems. This variant allows preemption early in request execution when memory consumption is low but restricts preemption as requests approach completion to optimize resource utilization. On the theoretical side, we derive a closed-form formula for this SRPT variant in an M/G/1 queue model, which demonstrates its potential value. In our system, we implement this preemption policy alongside our embedding-based prediction method. Our refined predictions from layer embeddings achieve 2.66x lower mean absolute error compared to BERT predictions from sequence prompts. TRAIL achieves 1.66x to 2.01x lower mean latency on the Alpaca dataset and 1.76x to 24.07x lower mean time to the first token compared to the state-of-the-art serving system.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a3105fa-2eb9-4fc3-92e6-15ee32e3f061Cited by top-tier papers8
- Fast Inference for Augmented Large Language ModelsRana Shahout, Cong Liang, Shiji Xin, Qianru Lao et al.NeurIPS 2025 · 13 citations
- Predicting LLM Output Length via Entropy-Guided RepresentationsHuanyi Xie, Yubin Chen, Liangyu Wang, Lijie Hu et al.ICLR 2026 · 12 citations
- Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM WorkloadsChaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi et al.NSDI 2026 · 5 citations
- AugServe: Adaptive Request Scheduling for Augmented Large Language Model Inference ServingYing Wang, Zhen Jin, Zhenqian Chen, Jiexiong Xu et al.ICML 2026 · 4 citations
- Scheduling LLM Inference with Uncertainty-Aware Output Length PredictionsHaoyu Zheng, Yongqiang Zhang, Fangcheng Fu, Xiaokai Zhou et al.ICML 2026 · 2 citations
Builds on5
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- Response Length Perception and Sequence Scheduling: An LLM-Empowered LLM Inference PipelineZangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo et al.NeurIPS 2023 · 159 citations
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas et al.HPCA 2025 · 106 citations
- SkipPredict: When to Invest in Predictions for SchedulingRana Shahout, Michael MitzenmacherNeurIPS 2024 · 6 citations
Related papers
- Efficient LLM Scheduling by Learning to RankYichao Fu, Siqi Zhu, Runlong Su, Aurick Qiao et al.NeurIPS 2024 · 129 citations
- Beyond Prediction: Tail-Aware Scheduling for LLM InferenceYueying Li, Yuanfan Chen, Jiayang Chen, Esha Choukse et al.ICML 2026
- FastServe: Iteration-Level Preemptive Scheduling for Large Language Model InferenceBingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu et al.NSDI 2026 · 12 citations
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae et al.HPDC 2026
- S3: Increasing GPU Utilization during Generative Inference for Higher ThroughputYunho Jin, Chun-Feng Wu, David Brooks, Gu-Yeon WeiNeurIPS 2023 · 150 citations
