PARD: Enhancing Goodput for Inference Pipeline via Proactive Request Dropping
Zhixin Zhao, Yitao Hu, Simin Chen, Mingfang Ji, Wei Yang, Yuhao Zhang, Laiping Zhao, Wenxin Li, Xiulong Liu, Wenyu Qu, Hao Wang
Abstract
Modern deep neural network (DNN) and large language model (LLM) applications integrate multiple models into inference pipelines with stringent latency requirements for customized tasks. To mitigate extensive request timeouts caused by accumulation, systems for inference pipelines commonly drop a subset of requests so the remaining ones can satisfy latency constraints. Since it is commonly believed that request dropping adversely affects goodput, existing systems only drop requests when they have to, which we call reactive dropping. However, this reactive policy can not maintain high goodput, as it neither makes timely dropping decisions nor identifies the proper set of requests to drop, leading to issues of dropping requests too late or dropping the wrong set of requests.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b338bbf-b4ef-4644-a9c2-e7d9cb9a22f2Builds on19
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Serving DNNs like Clockwork: Performance Predictability from the Bottom UpArpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao et al.OSDI 2020 · 392 citations
- Sinan: ML-based and QoS-aware resource management for cloud microservicesYanqi Zhang, Weizhe Hua, Zhuangzhuang Zhou, G. Edward Suh et al.ASPLOS 2021 · 226 citations
- Sage: practical and scalable ML-driven performance debugging in microservicesYu Gan, Mingyu Liang, Sundar Dev, David Lo et al.ASPLOS 2021 · 170 citations
Related papers
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning et al.NSDI 2026 · 29 citations
- Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsYufei Li, Zexin Li, Yinglun Zhu, Cong LiuRTSS 2025 · 4 citations
- SAGE: A Dataflow-Native Framework for Modular, Controllable, and Transparent LLM-Augmented ReasoningJun Liu, Peilin Liu, Ruicheng Zhang, Senlei Zhang et al.ICML 2026
- FlexPipe: Adapting Dynamic LLM Serving Through Inflight Pipeline Refactoring in Fragmented Serverless ClustersYanying Lin, Shijie Peng, Chengzhi Lu, ChengZhong Xu et al.EuroSys 2026 · 4 citations
- Tabi: An Efficient Multi-Level Inference System for Large Language ModelsYiding Wang, Kai Chen, Haisheng Tan, Kun GuoEuroSys 2023 · 64 citations
