ACBatch: Adaptive and Cooperative Batching for Edge Inference
Ziming Yang, Zichuan Zheng, Liyou Deng, Shan Zhang, Zhiyuan Wang, Hongbin Luo
摘要
Batching is a key technique in deep learning in-ference that enhances computational efficiency. Although widely applied in the cloud, batching may suffer from longer batch latency at edge servers due to highly dynamic task arrivals. In this paper, we propose an Adaptive and Cooperative Batching (ACBatch) framework for edge inference, wherein temporal adaptive batching and spatial task steering are jointly devised to balance the trade-off between batch latency and computational efficiency. To this end, a batch efficiency model is built to quantify the relationship between computational efficiency and batch size based on empirical measurements across diverse computing platforms and mainstream neural networks. Then, an optimization problem is formulated to minimize the completion time of a task sequence under ACBatch. For the simplified single-server case, the problem exhibits an optimal substructure and is solved by our proposed Dynamic Programming-based Adaptive Batching algorithm. For the general multi-server case, the optimization of ACBatch is proved NP-hard, and we propose the Multi-Server Cooperative Batching algorithm by iteratively optimizing batching and steering. Real-trace experiments show that ACBatch achieves an average improvement of 89.17% in completion time and 76.52% in latency compared to state-of-the-art methods.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- LSTCB: Long-Short-Timescale Cooperative Batching for Energy-Efficient Edge InferenceZichuan Zheng, Shan Zhang, Naixin Lu, Zhiyuan Wang 等INFOCOM 2026
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 被引用 184 次
- Joint Model and Data Adaptation for Cloud Inference ServingJingyan Jiang, Ziyue Luo, Chenghao Hu, Zhaoliang He 等RTSS 2021 · 被引用 19 次
- Towards Real-time Cooperative Deep Inference over the Cloud and Edge End DevicesShigeng Zhang, Yinggang Li, Xuan Liu, Song Guo 等UbiComp 2020 · 被引用 77 次
- Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning InferenceYujeong Choi, Yunseong Kim, Minsoo RhuHPCA 2021 · 被引用 65 次
