ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
Yuxing Xiang, Xue Li, Kun Qian, Yan Zhang, Wenyuan Yu, Ennan Zhai, Xin Jin, Jingren Zhou
Abstract
With the widespread adoption of Large Language Models (LLMs), serving LLM inference requests has become an increasingly important task, attracting active research advancements. Practical workloads play an essential role in this process: they are critical for motivating and benchmarking serving techniques and systems. However, the existing understanding of real-world LLM serving workloads is limited due to the lack of a comprehensive workload characterization. Prior analyses remain insufficient in scale and scope, thus failing to fully capture intricate workload characteristics.
In this paper, we fill the gap with an in-depth characterization of LLM serving workloads collected from our worldwide cloud LLM serving service, covering not only language models but also emerging multimodal and reasoning models, unveiling important new findings in each case. Moreover, based on our findings, we propose ServeGen, a principled framework for generating realistic LLM serving workloads by composing them on a per-client basis. Practical use cases validate that ServeGen achieves more accurate performance benchmarking compared to naive workload generation, and reveals new design implications that could otherwise be overlooked. ServeGen is open-sourced at https://github.com/alibaba/ServeGen.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory BallooningShan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li et al.OSDI 2026 · 33 citations
- Simple Is Better: Multiplication May Be All You Need for LLM Request SchedulingDingyan Zhang, Jinbo Han, Kaixi Zhang, Xingda Wei et al.OSDI 2026 · 5 citations
- SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM InferenceTian Xia, Ziming Mao, Jamison Kerney, Ethan J. Jackson et al.EuroSys 2026 · 2 citations
- Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the MarketYuxing Xiang, Xue Li, Kun Qian, Yufan Yang et al.SOSP 2025 · 2 citations
- NaturalSloth: Revisiting Denial-of-Service Attacks on Large Language ModelsYiming Chen, Zexin Li, Xianghu Yue, Robby T. Tan et al.ACL 2026
Builds on18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry et al.USENIX ATC 2020 · 946 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
Related papers
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 17 citations
- Batched in Back: Characterizing and Optimizing Offline LLM Inference in Production with ACDCLeping Yang, Xue Li, Kun Qian, Erci Xu et al.SOSP 2026
- KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud ProviderJiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen et al.USENIX ATC 2025 · 70 citations
- Large Language Models as Realistic Microservice Trace GeneratorsDonghyun Kim, Sriram Ravula, Taemin Ha, Alex Dimakis et al.EMNLP 2025 · 2 citations
- Reasoning Language Model Inference Serving Unveiled: An Empirical StudyQi Li, Junpan Wu, Xiang Liu, Yuxin Wang et al.ICLR 2026 · 3 citations
