Mil: Cost-guided Minimum Makespan Scheduling for Applications of Multiple LLMs
Jingzhi Fang, Yanyan Shen, Yue Wang, Lei Chen
摘要
Multi-LLM applications calling multiple LLMs per request are emerging. An important scenario is running these applications offline on a request set. This work aims to minimize the offline inference makespan of these applications to save time and cost. Specifically, we study minimum-makespan scheduling of multi-LLM applications (the MLAS problem), which requires GPU allocation, LLM parallelism selection, and LLM execution orchestration. MLAS is NP-hard, and it differs from existing multi-model frameworks and job scheduling problems due to LLMs' unique properties (e.g., high memory demand, complex inference behavior), the offline inference setting, and relaxed execution precedence constraints. There is no existing work on MLAS and simple rules cannot handle all the problem instances. We propose a framework, Mil, for MLAS with three major components: (1) processing functions estimating LLM processing rates by output length sampling, inference process simulation, and per-generation-iteration latency estimation; (2) a greedy method that finds a good schedule with a theoretical guarantee on a simplified problem instance; (3) a runtime adjustment mechanism reducing GPU idleness. Experiments on various applications (ensembling, routing, chain summary, mixed) show that Mil can achieve up to 3.4× end-to-end speedups over current practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- WizardLM: Empowering Large Pre-Trained Language Models to Follow Complex InstructionsCan Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng 等ICLR 2024 · 被引用 1,206 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
- RAPTOR: Recursive Abstractive Processing for Tree-Organized RetrievalParth Sarthi, Salman Abdullah, Aditi Tuli, Shubh Khanna 等ICLR 2024 · 被引用 460 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 被引用 173 次
相关 Paper
- Fast Inference for Augmented Large Language ModelsRana Shahout, Cong Liang, Shiji Xin, Qianru Lao 等NeurIPS 2025 · 被引用 13 次
- PKAS: Predictive KVCache-Aware Scheduling for Faster LLM and Transformer InferencesJie Ye, Avinash Maurya, Krishna Teja Chitty-Venkata, Bogdan Nicolae 等HPDC 2026
- Batched in Back: Characterizing and Optimizing Offline LLM Inference in Production with ACDCLeping Yang, Xue Li, Kun Qian, Erci Xu 等SOSP 2026
- Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsYufei Li, Zexin Li, Yinglun Zhu, Cong LiuRTSS 2025 · 被引用 4 次
- PiLLM: Resource-Efficient LLM Inference Using Workload PredictionYunqian Fan, Shihao Bai, Ruihao Gong, Zaijun Wang 等EuroSys 2026
