WIET: Harmonizing Group-aware Model Weighting and Worker Allocation for Ensemble Temporal Prediction MaaS
Binbin Feng, Shikun He, Yingxin Wang, Pengwei Wang, Xiang Gao, Zhijun Ding
Abstract
Ensemble Temporal Prediction Model-as-a-Service (ETP-MaaS) has become crucial in fields like financial modeling and cloud monitoring. Existing solutions fail to co-optimally address a two-fold challenge of dynamic collaboration and heterogeneity, treating models as independent entities and employing simplistic worker allocation rules. However, at the model level, data volatility means that optimal performance requires identifying and weighting constantly shifting subgroups of base models, not just individual ones; at the system level, these model groups must be efficiently mapped to a pool of heterogeneous and dynamically available workers. To this end, we introduce WIET, an efficient ETP-MaaS system that co-optimizes model weighting and worker allocation. For adaptive weighting, WIET identifies evolving group behaviors among base models and propose a novel group temporal locality-enhanced weighting method. Additionally, WIET develops an efficient, multi-dimensional worker allocation method powered by hybrid heuristic optimization, effectively reducing bottlenecks and resource waste. Experiments show WIET consistently outperforms state-of-the-art methods in terms of accuracy, latency, and resource usage across various workloads and tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on7
- Black-Box Tuning for Language-Model-as-a-ServiceTianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang et al.ICML 2022 · 343 citations
- INFless: a native serverless system for low-latency, high-throughput inferenceYanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang et al.ASPLOS 2022 · 145 citations
- Tabi: An Efficient Multi-Level Inference System for Large Language ModelsYiding Wang, Kai Chen, Haisheng Tan, Kun GuoEuroSys 2023 · 64 citations
- Towards Inference Efficient Deep Ensemble LearningZiyue Li, Kan Ren, Yifan Yang, Xinyang Jiang et al.AAAI 2023 · 18 citations
- Mastering Stock Markets with Efficient Mixture of Diversified Trading ExpertsShuo Sun, Xinrun Wang, Wanqi Xue, Xiaoxuan Lou et al.KDD 2023 · 13 citations
Related papers
- One for All: Unified Workload Prediction for Dynamic Multi-tenant Edge Cloud PlatformsShaoyuan Huang, Zheng Wang, Heng Zhang, Xiaofei Wang et al.KDD 2023 · 30 citations
- Cocktail: A Multidimensional Optimization for Model Serving in CloudJashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma et al.NSDI 2022
- Many Minds, One Goal: Time Series Forecasting via Sub-task Specialization and Inter-agent CooperationQihe Huang, Zhengyang Zhou, Yangze Li, Kuo Yang et al.NeurIPS 2025 · 11 citations
- Rethinking Cloud Optimization: Volatility-Driven for Better OutcomesBaoqing Wang, Gongming Zhao, Hongli Xu, Shibo Wu et al.SIGCOMM 2026
- PreServe: Intelligent Management for LMaaS Systems via Hierarchical PredictionZhihan Jiang, Yujie Huang, Guangba Yu, Junjie Huang et al.ICSE 2026 · 5 citations
