Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning Inference
Yujeong Choi, Yunseong Kim, Minsoo Rhu
Abstract
In cloud ML inference systems, batching is an essential technique to increase throughput which helps optimize total-cost-of-ownership. Prior graph batching combines the individual DNN graphs into a single one, allowing multiple inputs to be concurrently executed in parallel. We observe that the coarse-grained graph batching becomes suboptimal in effectively handling the dynamic inference request traffic, leaving significant performance left on the table. This paper proposes LazyBatching, an SLA-aware batching system that considers both scheduling and batching in the granularity of individual graph nodes, rather than the entire graph for flexible batching. We show that LazyBatching can intelligently determine the set of nodes that can be efficiently batched together, achieving an average 15×, 1.5×, and 5.5 × improvement than graph batching in terms of average response time, throughput, and SLA satisfaction, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal SharingSeungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park et al.USENIX ATC 2022 · 200 citations
- SHEPHERD: Serving DNNs in the WildHong Zhang, Yupeng Tang, Anurag Khandelwal, Ion StoicaNSDI 2023 · 161 citations
- OliVe: Accelerating Large Language Models via Hardware-friendly Outlier-Victim Pair QuantizationCong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng et al.ISCA 2023 · 151 citations
- ANT: Exploiting Adaptive Numerical Data Type for Low-bit Deep Neural Network QuantizationCong Guo, Chen Zhang, Jingwen Leng, Zihan Liu et al.MICRO 2022 · 109 citations
- Clover: Toward Sustainable AI with Carbon-Aware Machine Learning Inference ServiceBaolin Li, Siddharth Samsi, Vijay Gadepally, Devesh TiwariSC 2023 · 63 citations
Builds on3
- PREMA: A Predictive Multi-Task Scheduling Algorithm For Preemptible Neural Processing UnitsYujeong Choi, Minsoo RhuHPCA 2020 · 150 citations
- Centaur: A Chiplet-based, Hybrid Sparse-Dense Accelerator for Personalized RecommendationsRanggi Hwang, Taehun Kim, Youngeun Kwon, Minsoo RhuISCA 2020 · 94 citations
- NeuMMU: Architectural Support for Efficient Address Translations in Neural Processing UnitsBongjoon Hyun, Youngeun Kwon, Yujeong Choi, John Kim et al.ASPLOS 2020 · 29 citations
Related papers
- Harpagon: Minimizing DNN Serving Cost via Efficient Dispatching, Scheduling and SplittingZhixin Zhao, Yitao Hu, Ziqi Gong, Guotao Yang et al.INFOCOM 2025 · 2 citations
- DVABatch: Diversity-aware Multi-Entry Multi-Exit Batching for Efficient Processing of DNN Services on GPUsWeihao Cui, Han Zhao, Quan Chen, Hao Wei et al.USENIX ATC 2022 · 4 citations
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 184 citations
- Optimizing Inference Serving on Serverless PlatformsAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniVLDB 2022 · 76 citations
- ACBatch: Adaptive and Cooperative Batching for Edge InferenceZiming Yang, Zichuan Zheng, Liyou Deng, Shan Zhang et al.INFOCOM 2025 · 2 citations
