WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic Scheduling
Jingqi Feng, Yukai Huang, Rui Zhang, Sicheng Liang, Ming Yan, Jie Wu
Abstract
Existing large language model (LLM) serving systems typically batch the compute-bound prefill and I/O-bound decoding phases together.This co-location approach not only leads to significant interference between the two phases but also limits resource allocation and placements.To address these limitations, recent work proposes disaggregating the prefill and decoding phases to enhance performance.However, these works often rely on coarse-grained static scheduling strategies, resulting in imbalanced and insufficient resource utilization.For instance, compute resources for the prefill phase may be overloaded while those for the decoding phase remain idle, resulting in performance bottlenecks.In this paper, we propose WindServe, an efficient phase disaggregated LLM serving system that leverages stream-based, finegrained dynamic scheduling to enhance resource utilization and performance.WindServe features a global scheduler that monitors compute and memory resource usage to dynamically orchestrate cross-phase jobs, effectively reducing queuing delay and KV cache swapping overhead.We also introduce a stall-free rescheduling strategy to saturate the memory resources while minimizing the scheduling overhead from KV cache transfers.Furthermore, we design a stream-based approach to mitigate interference between prefill and decoding jobs.Our evaluation demonstrates that Wind-Serve achieves remarkable stability and SLO attainment under highload scenarios, outperforming state-of-the-art phase-disaggregated LLM serving systems by delivering a 4.28× improvement in TTFT median latency and a 1.5× reduction in TPOT P99 latency.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bf02c7dc-53a4-413f-a8dd-250c9646fa99Cited by top-tier papers7
- Towards High-Goodput LLM Serving with Prefill-decode MultiplexingYukang Chen, Weihao Cui, Han Zhao, Ziyi Xu et al.ASPLOS 2026 · 18 citations
- KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM ServingZedong Liu, Xinyang Ma, Dejun Luo, Hairui Zhao et al.SIGCOMM 2026 · 7 citations
- Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM WorkloadsChaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi et al.NSDI 2026 · 5 citations
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU MultiplexingLei Gao, Chaoyi Jiang, Hossein Entezari Zarch, Daniel Wong et al.ICML 2026 · 4 citations
- Weave: Efficient Co-Scheduling for Disaggregated RL Post-TrainingTianyuan Wu, Lunxi Cao, Yining Wei, Wei Gao et al.OSDI 2026
Related papers
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM ServingJiangfei Duan, Runyu Lu, Haojie Duanmu, Xiuhong Li et al.ICML 2024 · 51 citations
- LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence ParallelismBingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun et al.SOSP 2024 · 32 citations
- Efficient Multi-round LLM Inference over Disaggregated ServingWenhao He, Youhe Jiang, Penghao Zhao, Quanqing Xu et al.ICML 2026 · 7 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
