SCOOT: SLO-Oriented Performance Tuning for LLM Inference Engines
Ke Cheng, Zhi Wang, Wen Hu, Tiannuo Yang, Jianguo Li, Sheng Zhang
摘要
As large language models (LLMs) are gaining increasing popularity across a wide range of web applications, it is of great importance to optimize service-level objectives (SLOs) for LLM inference services to enhance user satisfaction and improve the competitiveness of cloud vendors. In this paper, we observe that adjusting the parameters of LLM inference engines can improve service performance, and the optimal parameter configurations of different services are different. Therefore, we propose SCOOT, an automatic performance tuning system to optimize SLOs for each LLM inference service by tuning the parameters of the inference engine. SCOOT jointly exploits single-objective and multiple-objective Bayesian optimization (BO) techniques to handle various optimization objectives via exploration and exploitation. Moreover, SCOOT prunes the search space with known constraints and adopts a random forest to learn hidden constraints during the tuning process to mitigate invalid exploration. To improve the tuning efficiency, SCOOT utilizes the parallel suggestion to accelerate the tuning process. Extensive experiments demonstrate that SCOOT considerably outperforms existing tuning techniques in SLO optimization while greatly improving the tuning efficiency. Moreover, SCOOT is universally applicable to various LLM inference engines including vLLM and TensorRT-LLM. Currently, SCOOT has already been implemented in the production environment at Ant Group. CCS Concepts • Computing methodologies → Parallel computing methodologies; Machine learning; Natural language generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Bullet: Boosting GPU Utilization for LLM Serving via Dynamic Spatial-Temporal OrchestrationZejia Lin, Hongxin Xu, Guanyi Chen, Zhiguang Chen 等ASPLOS 2026 · 被引用 2 次
- InferLog: Accelerating LLM Inference for Online Log Parsing via ICL-oriented Prefix CachingYilun Wang, Pengfei Chen, Haiyu Huang, Zilong He 等ICSE 2026 · 被引用 1 次
- Nova: Real-Time Agentic Vision-Language Model Serving With Adaptive Cross-Stage ParallelizationYuhang Xu, Shengzhong Liu, Dong Zhang, Bingheng Yan 等RTSS 2025 · 被引用 1 次
- Experiential Fairness: Bridging the Gap Between User Experience and Resource-Centric Fairness in Online LLM ServicesJiahua Huang, Wentai Wu, Yongheng Liu, Guozhi Liu 等AAAI 2026
它引用的顶会 Paper14
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
相关 Paper
- VDTuner: Automated Performance Tuning for Vector Data Management SystemsTiannuo Yang, Wen Hu, Wangqi Peng, Yusen Li 等ICDE 2024 · 被引用 12 次
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 被引用 17 次
- Resource Multiplexing in Tuning and Serving Large Language ModelsYongjun He, Haofeng Yang, Yao Lu, Ana Klimovic 等USENIX ATC 2025 · 被引用 9 次
- Inference-Aware Prompt Optimization for Aligning Black-Box Large Language ModelsSaaduddin Mahmud, Mason Nakamura, Kyle Hollins Wray, Shlomo ZilbersteinAAAI 2026
- λ-Tune: Harnessing Large Language Models for Automated Database System TuningVictor Giannakouris, Immanuel TrummerSIGMOD 2025 · 被引用 20 次
