TurboTransformers: an efficient GPU serving system for transformer models
Jiarui Fang, Yang Yu, Chengduo Zhao, Jie Zhou
Abstract
The transformer is the most critical algorithm innovation of the Nature Language Processing (NLP) field in recent years. Unlike the Recurrent Neural Network (RNN) models, Transformers can process on dimensions of sequence lengths in parallel, therefore leading to better accuracy on long sequences. However, efficient deployments of them for online services in data centers equipped with GPUs are not easy. First, more computation introduced by transformer structures makes it more challenging to meet the latency and throughput constraints of serving. Second, NLP tasks take in sentences of variable length. The variability of input dimensions brings a severe problem to efficient memory management and serving optimization.
This paper designed a transformer serving system called TurboTransformers, which consists of a computing runtime and a serving framework to solve the above challenges. Three innovative features make it stand out from other similar works. An efficient parallel algorithm is proposed for GPU-based batch reduction operations, like Softmax and LayerNorm, major hot spots besides BLAS routines. A memory allocation algorithm, which better balances the memory footprint and allocation/free efficiency, is designed for variable-length input situations. A serving framework equipped with a new batch scheduler using dynamic programming achieves the optimal throughput on variable-length requests. The system can achieve the state-of-the-art transformer model serving performance on GPU platforms and can be seamlessly integrated into your PyTorch code with a few lines of code.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8042987-af42-4304-85f4-07c3ab09375fCited by top-tier papers59
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim et al.OSDI 2022 · 690 citations
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li et al.ICML 2023 · 683 citations
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model ServingYinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu et al.OSDI 2024 · 646 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
Builds on1
Related papers
- PetS: A Unified Framework for Parameter-Efficient Transformers ServingZhe Zhou, Xuechao Wei, Jiejing Zhang, Guangyu SunUSENIX ATC 2022 · 41 citations
- A length adaptive algorithm-hardware co-design of transformer on FPGA through sparse attention and dynamic pipeliningHongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li et al.DAC 2022 · 49 citations
- DeepSpeed- Inference: Enabling Efficient Inference of Transformer Models at Unprecedented ScaleReza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li et al.SC 2022 · 276 citations
- ALISA: Accelerating Large Language Model Inference via Sparsity-Aware KV CachingYoupeng Zhao, Di Wu, Jun WangISCA 2024 · 35 citations
- OTAS: An Elastic Transformer Serving System via Token AdaptationJinyu Chen, Wenchao Xu, Zicong Hong, Song Guo et al.INFOCOM 2024 · 4 citations
