A Resource-centric Analysis and Optimization of NoSQL Workloads using Distressed Resource Volume Metric
Gunika Verma, Aashutosh A V, Pooja Srinivas, Yogesh Simmhan, Ayush Choure, Harshit Shah, Mayukh Das, Prashant Sasatte, Chetan Bansal, Abhijit Pai, Suraj Dixit, Achint Agrawal
摘要
Large-scale managed cloud databases leverage sophisticated load Packing and Migration (PAM) algorithms, which provide the efficiencies necessary for running these services at scale on cloud resources. Research into optimizing the resources and reliability of cloud databases at massive scales is limited by a lack of public NoSQL workloads. We address this in the context of Cosmos DB , Microsoft's flagship cloud-hosted NoSQL database. We first propose open-source NoSQL workloads from real Cosmos DB clusters, and analyze these traces to derive a novel reliability metric, Distressed Resource Volume (DRV) , which captures the quality of service experienced by the end user. We then develop an open-source policy simulation framework, LoadStar , powered by a non-parametric statistical model of estimating the QoS of real traffic patterns. These form a reusable benchmark pipeline for validating policies for resource-centric NoSQL workloads. We then define a resource optimization problem for placing Cosmos DB replicas onto VM nodes, develop the Luna model for forecasting future load distributions, and the Orbit PAM algorithm that uses these forecasts to trigger and rebalance stressed replicas, to reduce tail-errors. Our experiments, validated using LoadStar for these workloads, demonstrate Orbit's benefits over the existing Cosmos DB policy and a worst-fit optimized baseline, with higher load delivered at lower error rates and up to 35% reduction in resources. These have been deployed in production, with potential savings of $100 Ms /yr while improving service reliability for millions of customers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- ResTune: Resource Oriented Tuning Boosted by Meta-Learning for Cloud DatabasesXinyi Zhang, Hong Wu, Zhuo Chang, Shuowei Jin 等SIGMOD 2021 · 被引用 113 次
- OPTIMUSCLOUD: Heterogeneous Configuration Optimization for Distributed Databases in the CloudAshraf Mahgoub, Alexander Medoff, Rakesh Kumar, Subrata Mitra 等USENIX ATC 2020 · 被引用 63 次
- Moneyball: Proactive Auto-Scaling in Microsoft Azure SQL Database ServerlessOlga Poppe, Qun Guo, Willis Lang, Pankaj Arora 等VLDB 2022 · 被引用 40 次
- TSM-Bench: Benchmarking Time Series Database Systems for Monitoring ApplicationsAbdelouahab Khelifati, Mourad Khayati, Anton Dignös, Djellel Eddine Difallah 等VLDB 2023 · 被引用 24 次
- Don't Look Back, Look into the Future: Prescient Data Partitioning and Migration for Deterministic Database SystemsYu-Shan Lin, Ching Tsai, Tz-Yu Lin, Yun-Sheng Chang 等SIGMOD 2021 · 被引用 17 次
相关 Paper
- Tenant Placement in Over-subscribed Database-as-a-Service ClustersArnd Christian König, Yi Shan, Tobias Ziegler, Aarati Kakaraparthy 等VLDB 2022 · 被引用 8 次
- Flexible Resource Allocation for Relational Database-as-a-ServicePankaj Arora, Surajit Chaudhuri, Sudipto Das, Junfeng Dong 等VLDB 2023 · 被引用 10 次
- Runtime Variation in Big Data AnalyticsYiwen Zhu, Rathijit Sen, Robert Horton, John Mark AgostaSIGMOD 2023 · 被引用 5 次
- CloudyBench: A Testbed for A Comprehensive Evaluation of Cloud-Native DatabasesChao Zhang, Guoliang Li, Leyao Liu, Tao Lv 等ICDE 2025 · 被引用 4 次
- Seagull: An Infrastructure for Load Prediction and Optimized Resource AllocationOlga Poppe, Tayo Amuneke, Dalitso Banda, Aritra De 等VLDB 2021 · 被引用 37 次
