LSTCB: Long-Short-Timescale Cooperative Batching for Energy-Efficient Edge Inference
Zichuan Zheng, Shan Zhang, Naixin Lu, Zhiyuan Wang, Hongbin Luo
Abstract
Batching is a critical technique for improving energy efficiency (EE) when dealing with edge-based inference tasks. While existing batching policies achieve significant energy savings by exploiting short-timescale arrival patterns, they largely overlook long-timescale load fluctuations—leading to severe EE degradation during lightly-loaded periods, as confirmed by prototype measurements. To overcome this limitation, we propose Long-Short-Timescale Cooperative Batching (LSTCB), a novel framework that enhances EE without sacrificing latency. LSTCB dynamically deactivates servers in response to load variations while intelligently coordinating batched inference tasks. To achieve this, we first develop a measurement-based EE model for edge servers that accurately characterizes the relationship among batch size, workload, and energy consumption. At the short timescale, we devise a dynamic batching policy through Lyapunov optimization that simultaneously minimizes both energy consumption and latency under varying traffic conditions. For long-timescale optimization, we implement an intelligent server deactivation scheme that employs dynamic programming combined with bipartite graph matching to steer inter-server traffic, thereby minimizing overall system-level energy costs while maintaining service quality. Real-world prototype experiments demonstrate that LSTCB achieves 24.48%-46.12% energy savings over baseline methods while maintaining competitive latency performance.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 6ea02ecd-fec0-47ba-87df-84192585be39Related papers
- ACBatch: Adaptive and Cooperative Batching for Edge InferenceZiming Yang, Zichuan Zheng, Liyou Deng, Shan Zhang et al.INFOCOM 2025 · 2 citations
- Optimizing Inference Serving on Serverless PlatformsAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniVLDB 2022 · 76 citations
- Many Hands Make Light Work: Accelerating Edge Inference via Multi-Client Collaborative CachingWenyi Liang, Jianchun Liu, Hongli Xu, Chunming Qiao et al.ICDE 2025
- Batch: machine learning inference serving on serverless platforms with adaptive batchingAhsan Ali, Riccardo Pinciroli, Feng Yan, Evgenia SmirniSC 2020 · 184 citations
- Lazy Batching: An SLA-aware Batching System for Cloud Machine Learning InferenceYujeong Choi, Yunseong Kim, Minsoo RhuHPCA 2021 · 65 citations
