PASS: Predictive Auto-Scaling System for Large-scale Enterprise Web Applications
Yunda Guo, Jiake Ge, Panfeng Guo, Yunpeng Chai, Tao Li, Mengnan Shi, Yang Tu, Jian Ouyang
Abstract
We confront two challenges in the management of a vast and diverse array of online web applications deployed on enterprise-grade autoscaling infrastructure, primarily focused on ensuring Quality of Service (QoS) for large-scale applications and optimizing resource costs. Firstly, reacting to increased load with a response-based approach can temporarily degrade QoS because many web applications need a few minutes to warm up. Therefore, precise workload prediction is critical for predictive scaling. However, our analysis of real-world applications underscores the substantial challenges arising from the limited precision and robustness of existing single prediction algorithms in the context of predictive auto-scaling. Secondly, guaranteeing the QoS of online applications within a costeffective structure is crucial, as it is inherently linked to corporate profitability. Nevertheless, our study shows that mainstream autoscaling methods exhibit various limitations, either being unsuitable for online environments or inadequately ensuring QoS. To address these issues, we introduce PASS, a Predictive Auto-Scaling System tailored for large-scale online web applications in enterprise settings. Our highly robust and accurate prediction framework dynamically integrates and calibrates appropriate prediction algorithms based on the unique characteristics of each application to effectively manage workload diversity. We further establish a performance model derived from online historical logs, enhancing auto-scaling to ensure diverse QoS without adverse impacts on online applications. Additionally, we implement a reactive strategy grounded in queuing theory to promptly address QoS violations resulting from inaccurate predictions or unexpected events. Across a wide spectrum of applications and real-world workloads, PASS outperforms state-of-the-art methods, achieving higher workload prediction accuracy and a superior QoS guarantee rate with less resource cost. KEYWORDS auto-scaling, workload prediction, quality of service, performance model, cloud computing 1 the mutation features are especially prevalent and critical in enterprise-level applications. This is due to a substantial influx of QPS during morning, noon, and evening peak hours, leading to a considerable and undeniable impact.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Exploiting Language Power for Time Series Forecasting with Exogenous VariablesQihe Huang, Zhengyang Zhou, Kuo Yang, Yang WangWWW 2025 · 17 citations
- PreServe: Intelligent Management for LMaaS Systems via Hierarchical PredictionZhihan Jiang, Yujie Huang, Guangba Yu, Junjie Huang et al.ICSE 2026 · 5 citations
Builds on9
- TS2Vec: Towards Universal Representation of Time SeriesZhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang et al.AAAI 2022 · 938 citations
- Self-Supervised Contrastive Pre-Training For Time Series via Time-Frequency ConsistencyXiang Zhang, Ziyuan Zhao, Theodoros Tsiligkaridis, Marinka ZitnikNeurIPS 2022 · 558 citations
- A Time Series is Worth 64 Words: Long-term Forecasting with TransformersYuqi Nie, Nam H. Nguyen, Phanwadee Sinthong, Jayant KalagnanamICLR 2023 · 536 citations
- Neural Transformation Learning for Deep Anomaly Detection Beyond ImagesChen Qiu, Timo Pfrommer, Marius Kloft, Stephan Mandt et al.ICML 2021 · 171 citations
- Omni-Scale CNNs: a simple and effective kernel size configuration for time series classificationWensi Tang, Guodong Long, Lu Liu, Tianyi Zhou et al.ICLR 2022 · 163 citations
Related papers
- OffDQ: An Offline Deep Learning Framework for QoS PredictionSoumi Chattopadhyay, Richik Chanda, Suraj Kumar, Chandranath AdakWWW 2022 · 11 citations
- RobustScaler: QoS-Aware Autoscaling for Complex WorkloadsHuajie Qian, Qingsong Wen, Liang Sun, Jing Gu et al.ICDE 2022 · 41 citations
- The Fast and The Frugal: Tail Latency Aware Provisioning for Coping with Load VariationsAdithya Kumar, Iyswarya Narayanan, Timothy Zhu, Anand SivasubramaniamWWW 2020 · 20 citations
- Robust Auto-Scaling with Probabilistic Workload Forecasting for Cloud DatabasesHaitian Hang, Xiu Tang, Jianling Sun, Lingfeng Bao et al.ICDE 2024 · 11 citations
- Predicting Failures of Autoscaling Distributed ApplicationsGiovanni Denaro, Noura El Moussa, Rahim Heydarov, Francesco Lomio et al.FSE 2024 · 2 citations
