WindServe: Efficient Phase-Disaggregated LLM Serving with Stream-based Dynamic Scheduling
Jingqi Feng, Yukai Huang, Rui Zhang, Sicheng Liang, Ming Yan, Jie Wu
摘要
Existing large language model (LLM) serving systems typically batch the compute-bound prefill and I/O-bound decoding phases together.This co-location approach not only leads to significant interference between the two phases but also limits resource allocation and placements.To address these limitations, recent work proposes disaggregating the prefill and decoding phases to enhance performance.However, these works often rely on coarse-grained static scheduling strategies, resulting in imbalanced and insufficient resource utilization.For instance, compute resources for the prefill phase may be overloaded while those for the decoding phase remain idle, resulting in performance bottlenecks.In this paper, we propose WindServe, an efficient phase disaggregated LLM serving system that leverages stream-based, finegrained dynamic scheduling to enhance resource utilization and performance.WindServe features a global scheduler that monitors compute and memory resource usage to dynamically orchestrate cross-phase jobs, effectively reducing queuing delay and KV cache swapping overhead.We also introduce a stall-free rescheduling strategy to saturate the memory resources while minimizing the scheduling overhead from KV cache transfers.Furthermore, we design a stream-based approach to mitigate interference between prefill and decoding jobs.Our evaluation demonstrates that Wind-Serve achieves remarkable stability and SLO attainment under highload scenarios, outperforming state-of-the-art phase-disaggregated LLM serving systems by delivering a 4.28× improvement in TTFT median latency and a 1.5× reduction in TPOT P99 latency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper7
- Towards High-Goodput LLM Serving with Prefill-decode MultiplexingYukang Chen, Weihao Cui, Han Zhao, Ziyi Xu 等ASPLOS 2026 · 被引用 18 次
- KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM ServingZedong Liu, Xinyang Ma, Dejun Luo, Hairui Zhao 等SIGCOMM 2026 · 被引用 7 次
- Libra: Flexible Request Partitioning and Scheduling for Serving Unbalanced and Dynamic LLM WorkloadsChaoyi Ruan, Yinhe Chen, Dongqi Tian, Yandong Shi 等NSDI 2026 · 被引用 5 次
- DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU MultiplexingLei Gao, Chaoyi Jiang, Hossein Entezari Zarch, Daniel Wong 等ICML 2026 · 被引用 4 次
- Weave: Efficient Co-Scheduling for Disaggregated RL Post-TrainingTianyuan Wu, Lunxi Cao, Yining Wei, Wei Gao 等OSDI 2026
相关 Paper
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu 等OSDI 2024 · 被引用 646 次
- MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM ServingJiangfei Duan, Runyu Lu, Haojie Duanmu, Xiuhong Li 等ICML 2024 · 被引用 51 次
- LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence ParallelismBingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun 等SOSP 2024 · 被引用 32 次
- Efficient Multi-round LLM Inference over Disaggregated ServingWenhao He, Youhe Jiang, Penghao Zhao, Quanqing Xu 等ICML 2026 · 被引用 7 次
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan 等OSDI 2024 · 被引用 537 次
