ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
Yuxing Xiang, Xue Li, Kun Qian, Yan Zhang, Wenyuan Yu, Ennan Zhai, Xin Jin, Jingren Zhou
摘要
With the widespread adoption of Large Language Models (LLMs), serving LLM inference requests has become an increasingly important task, attracting active research advancements. Practical workloads play an essential role in this process: they are critical for motivating and benchmarking serving techniques and systems. However, the existing understanding of real-world LLM serving workloads is limited due to the lack of a comprehensive workload characterization. Prior analyses remain insufficient in scale and scope, thus failing to fully capture intricate workload characteristics.
In this paper, we fill the gap with an in-depth characterization of LLM serving workloads collected from our worldwide cloud LLM serving service, covering not only language models but also emerging multimodal and reasoning models, unveiling important new findings in each case. Moreover, based on our findings, we propose ServeGen, a principled framework for generating realistic LLM serving workloads by composing them on a per-client basis. Practical use cases validate that ServeGen achieves more accurate performance benchmarking compared to naive workload generation, and reveals new design implications that could otherwise be overlooked. ServeGen is open-sourced at https://github.com/alibaba/ServeGen.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Prism: Cost-Efficient Multi-LLM Serving via GPU Memory BallooningShan Yu, Yifan Qiao, Mingyuan Ma, Yangmin Li 等OSDI 2026 · 被引用 33 次
- Simple Is Better: Multiplication May Be All You Need for LLM Request SchedulingDingyan Zhang, Jinbo Han, Kaixi Zhang, Xingda Wei 等OSDI 2026 · 被引用 5 次
- SkyWalker: A Locality-Aware Cross-Region Load Balancer for LLM InferenceTian Xia, Ziming Mao, Jamison Kerney, Ethan J. Jackson 等EuroSys 2026 · 被引用 2 次
- Aegaeon: Effective GPU Pooling for Concurrent LLM Serving on the MarketYuxing Xiang, Xue Li, Kun Qian, Yufan Yang 等SOSP 2025 · 被引用 2 次
- NaturalSloth: Revisiting Denial-of-Service Attacks on Large Language ModelsYiming Chen, Zexin Li, Xianghu Yue, Robby T. Tan 等ACL 2026
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun 等NeurIPS 2024 · 被引用 1,586 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud ProviderMohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Irfan Chaudhry 等USENIX ATC 2020 · 被引用 946 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
相关 Paper
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 被引用 17 次
- Batched in Back: Characterizing and Optimizing Offline LLM Inference in Production with ACDCLeping Yang, Xue Li, Kun Qian, Erci Xu 等SOSP 2026
- KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud ProviderJiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen 等USENIX ATC 2025 · 被引用 70 次
- Large Language Models as Realistic Microservice Trace GeneratorsDonghyun Kim, Sriram Ravula, Taemin Ha, Alex Dimakis 等EMNLP 2025 · 被引用 2 次
- Reasoning Language Model Inference Serving Unveiled: An Empirical StudyQi Li, Junpan Wu, Xiang Liu, Yuxin Wang 等ICLR 2026 · 被引用 3 次
