LSTCB: Long-Short-Timescale Cooperative Batching for Energy-Efficient Edge Inference
Zichuan Zheng, Shan Zhang, Naixin Lu, Zhiyuan Wang, Hongbin Luo
摘要
Batching is a critical technique for improving energy efficiency (EE) when dealing with edge-based inference tasks. While existing batching policies achieve significant energy savings by exploiting short-timescale arrival patterns, they largely overlook long-timescale load fluctuations—leading to severe EE degradation during lightly-loaded periods, as confirmed by prototype measurements. To overcome this limitation, we propose Long-Short-Timescale Cooperative Batching (LSTCB), a novel framework that enhances EE without sacrificing latency. LSTCB dynamically deactivates servers in response to load variations while intelligently coordinating batched inference tasks. To achieve this, we first develop a measurement-based EE model for edge servers that accurately characterizes the relationship among batch size, workload, and energy consumption. At the short timescale, we devise a dynamic batching policy through Lyapunov optimization that simultaneously minimizes both energy consumption and latency under varying traffic conditions. For long-timescale optimization, we implement an intelligent server deactivation scheme that employs dynamic programming combined with bipartite graph matching to steer inter-server traffic, thereby minimizing overall system-level energy costs while maintaining service quality. Real-world prototype experiments demonstrate that LSTCB achieves 24.48%-46.12% energy savings over baseline methods while maintaining competitive latency performance.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- ACBatch: Adaptive and Cooperative Batching for Edge InferenceZiming Yang, Zichuan Zheng, Liyou Deng, Shan Zhang 等INFOCOM 2025 · 被引用 2 次
- Optimizing Inference Serving on Serverless PlatformsAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniVLDB 2022 · 被引用 76 次
- Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative CachingWenyi Liang, Jianchun Liu, Hongli Xu, Chunming Qiao 等ICDE 2025
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 被引用 184 次
- Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning InferenceYujeong Choi, Yunseong Kim, Minsoo RhuHPCA 2021 · 被引用 65 次
