XY-Serve: End-to-End Versatile Production Serving for Dynamic LLM Workloads
Mingcong Song, Xinru Tang, Fengfan Hou, Jing Li, Wei Wei, Yipeng Ma, Runqiu Xiao, Hongjie Si, Dingcheng Jiang, Shouyi Yin, Yang Hu, Guoping Long
摘要
Meeting growing demands for low latency and cost efficiency in production-grade large language model (LLM) serving systems requires integrating advanced optimization techniques. However, dynamic and unpredictable input-output lengths of LLM, compounded by these optimizations, exacerbate the issues of workload variability, making it difficult to maintain high efficiency on AI accelerators, especially DSAs with tile-based programming models. To address this challenge, we introduce XY-Serve, a versatile, Ascend NPU native, end-to-end production LLM-serving system. The core idea is an abstraction mechanism that smooths out the workload variability by decomposing computations into unified, hardware-friendly, fine-grained meta primitives. Then, kernels can efficiently execute without concerning the irregularity of workload. After this abstraction mechanism, for Attention, we propose a meta-kernel that computes the basic pattern of GEMM-Softmax-GEMM with architectural-aware tile sizes. For Linear, we introduce a virtual padding scheme that adapts to dynamic shape changes while using highly efficient GEMM primitives with assorted fixed tile sizes. XY-Serve sits harmoniously with vLLM. Experimental results show up to 95% end-to-end throughput improvement compared with current publicly available baselines on Ascend NPUs. We also set a new performance record for Linear (average 14.6% faster) and Attention (average 21.5% faster) kernels relative to existing libraries. Lastly, we demonstrate the generality of our technologies on GPU platform.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- DEEPSERVE: Serverless Large Language Model Serving at ScaleJunhao Hu, Jiang Xu, Zhixia Liu, Yulong He 等USENIX ATC 2025 · 被引用 38 次
- QiMeng-Xpiler: Transcompiling Tensor Programs for Deep Learning Systems with a Neural-Symbolic ApproachShouyang Dong, Jun Bi, Di Huang, Jiaming Guo 等OSDI 2025 · 被引用 6 次
- SpaceServe: Spatial Multiplexing of Complementary Encoders and Decoders for Multimodal LLMsZhicheng Li, Shuoming Zhang, Jiacheng Zhao, Siqi Li 等NeurIPS 2025 · 被引用 5 次
- HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End OptimizationSize Li, Zhiqing Tang, Hongrui Liang, Jianxiong Guo 等ICML 2026
- NanoFlow: Towards Optimal Large Language Model Serving ThroughputKan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao 等OSDI 2025 · 被引用 92 次
