WIET: Harmonizing Group-aware Model Weighting and Worker Allocation for Ensemble Temporal Prediction MaaS
Binbin Feng, Shikun He, Yingxin Wang, Pengwei Wang, Xiang Gao, Zhijun Ding
摘要
Ensemble Temporal Prediction Model-as-a-Service (ETP-MaaS) has become crucial in fields like financial modeling and cloud monitoring. Existing solutions fail to co-optimally address a two-fold challenge of dynamic collaboration and heterogeneity, treating models as independent entities and employing simplistic worker allocation rules. However, at the model level, data volatility means that optimal performance requires identifying and weighting constantly shifting subgroups of base models, not just individual ones; at the system level, these model groups must be efficiently mapped to a pool of heterogeneous and dynamically available workers. To this end, we introduce WIET, an efficient ETP-MaaS system that co-optimizes model weighting and worker allocation. For adaptive weighting, WIET identifies evolving group behaviors among base models and propose a novel group temporal locality-enhanced weighting method. Additionally, WIET develops an efficient, multi-dimensional worker allocation method powered by hybrid heuristic optimization, effectively reducing bottlenecks and resource waste. Experiments show WIET consistently outperforms state-of-the-art methods in terms of accuracy, latency, and resource usage across various workloads and tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Black-Box Tuning for Language-Model-as-a-ServiceTianxiang Sun, Yunfan Shao, Hong Qian, Xuanjing Huang 等ICML 2022 · 被引用 343 次
- INFless: a native serverless system for low-latency, high-throughput inferenceYanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang 等ASPLOS 2022 · 被引用 145 次
- Tabi: An Efficient Multi-Level Inference System for Large Language ModelsYiding Wang, Kai Chen, Haisheng Tan, Kun GuoEuroSys 2023 · 被引用 64 次
- Towards Inference Efficient Deep Ensemble LearningZiyue Li, Kan Ren, Yifan Yang, Xinyang Jiang 等AAAI 2023 · 被引用 18 次
- Mastering Stock Markets with Efficient Mixture of Diversified Trading ExpertsShuo Sun, Xinrun Wang, Wanqi Xue, Xiaoxuan Lou 等KDD 2023 · 被引用 13 次
相关 Paper
- One for All: Unified Workload Prediction for Dynamic Multi-tenant Edge Cloud PlatformsShaoyuan Huang, Zheng Wang, Heng Zhang, Xiaofei Wang 等KDD 2023 · 被引用 30 次
- Cocktail: A Multidimensional Optimization for Model Serving in CloudJashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma 等NSDI 2022
- Many Minds, One Goal: Time Series Forecasting via Sub-task Specialization and Inter-agent CooperationQihe Huang, Zhengyang Zhou, Yangze Li, Kuo Yang 等NeurIPS 2025 · 被引用 11 次
- Rethinking Cloud Optimization: Volatility-Driven for Better OutcomesBaoqing Wang, Gongming Zhao, Hongli Xu, Shibo Wu 等SIGCOMM 2026
- PreServe: Intelligent Management for LMaaS Systems via Hierarchical PredictionZhihan Jiang, Yujie Huang, Guangba Yu, Junjie Huang 等ICSE 2026 · 被引用 5 次
