Batched in Back: Characterizing and Optimizing Offline LLM Inference in Production with ACDC
Leping Yang, Xue Li, Kun Qian, Erci Xu, Mingzhen Han, Haoran Zhu, Tao He, Zuolong Yin, Ennan Zhai, Wenyuan Yu, Jingren Zhou, Guangtao Xue
摘要
Serving offline large language model (LLM) inference workloads (e.g., log summarization and bulk translation) can consume up to 30% of GPUs in production. Despite this significant share, the characteristics of offline inference remain largely understudied. In this paper, we start by analyzing 1.5 million tasks comprising 23 billion requests across text and multi-modal models. We discover that the key properties of offline workloads, namely inherent determinism and throughput orientation, are neither exploited by online LLM serving systems nor by existing offline serving frameworks.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- PiLLM: Resource-Efficient LLM Inference Using Workload PredictionYunqian Fan, Shihao Bai, Ruihao Gong, Zaijun Wang 等EuroSys 2026
- ServeGen: Workload Characterization and Generation of Large Language Model Serving in ProductionYuxing Xiang, Xue Li, Kun Qian, Yan Zhang 等NSDI 2026 · 被引用 58 次
- ConServe: Fine-Grained GPU Harvesting for LLM Online and Offline Co-ServingYifan Qiao, Shan Yu, Shu Anzai, Haoran Ma 等ICML 2026 · 被引用 7 次
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 被引用 17 次
- Lemix: Unified Scheduling for Llm Training and Inference on Multi-Gpu SystemsYufei Li, Zexin Li, Yinglun Zhu, Cong LiuRTSS 2025 · 被引用 4 次
