On-Demand State Separation for Cloud Data Warehousing
Christian Winter, Jana Giceva, Thomas Neumann, Alfons Kemper
摘要
Moving data analysis and processing to the cloud is no longer reserved for a few companies with petabytes of data. Instead, the flexibility of on-demand resources is attracting an increasing number of customers with small to medium-sized workloads. These workloads do not occupy entire clusters but can run on single worker machines. However, picking the right worker for the job is challenging. Abstracting from worker machines, e.g., using stateless architectures, introduces overheads impacting performance. Solutions without stateless architectures resort to query restarts in the event of an adverse worker matching, wasting already achieved progress. In this paper, we propose migrating queries between workers by introducing on-demand state separation. Using state separation only when required enables maximum flexibility and performance while keeping already achieved progress. To derive the requirements for state separation, we first analyze the query state of medium-sized workloads on the example of TPC-DS SF100. Using this, we analyze the cost and describe the constraints necessary for state separation on such a workload. Furthermore, we describe the design and implementation of on-demand state separation in a compiling database system. Finally, using this implementation, we show the feasibility of our approach on TPC-DS and give a detailed analysis of the cost of query migration and state separation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 被引用 45 次
- CXL Memory Performance for In-Memory Data ProcessingMarcel Weisgut, Daniel Ritter, Pinar Tözün, Lawrence Benson 等VLDB 2025 · 被引用 7 次
- Towards Designing Future-Proof Data Processing SystemsMichael Jungmair, Jana GicevaVLDB 2025 · 被引用 1 次
- Interoperable ACID Transactions for Open Table FormatsTobias Götz, Daniel Ritter, Muhammad El-Hindi, Jana GicevaVLDB 2026
它引用的顶会 Paper6
- Building An Elastic Query Engine on Disaggregated StorageMidhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong 等NSDI 2020 · 被引用 142 次
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu 等VLDB 2021 · 被引用 67 次
- Towards Cost-Optimal Query Processing in the CloudViktor Leis, Maximilian KuschewskiVLDB 2021 · 被引用 34 次
- Self-Tuning Query Scheduling for Analytical WorkloadsBenjamin Wagner, André Kohn, Thomas NeumannSIGMOD 2021 · 被引用 25 次
- Crystal: A Unified Cache Storage System for Analytical DatabasesDominik Durner, Badrish Chandramouli, Yinan LiVLDB 2021 · 被引用 11 次
相关 Paper
- Understanding the Effect of Data Center Resource Disaggregation on Production DBMSsQizhen Zhang, Yifan Cai, Xinyi Chen, Sebastian Angel 等VLDB 2020 · 被引用 64 次
- OPTIMUSCLOUD: Heterogeneous Configuration Optimization for Distributed Databases in the CloudAshraf Mahgoub, Alexander Medoff, Rakesh Kumar, Subrata Mitra 等USENIX ATC 2020 · 被引用 63 次
- Cloud Analytics BenchmarkAlexander van Renen, Viktor LeisVLDB 2023 · 被引用 32 次
- Rhino: Efficient Management of Very Large Distributed State for Stream Processing EnginesBonaventura Del Monte, Steffen Zeuch, Tilmann Rabl, Volker MarklSIGMOD 2020 · 被引用 56 次
- AQD: Online Adaptive Query Dispatcher for HTAP DatabasesYang Wu, Tongliang Li, Xuanhe Zhou, Jianying Wang 等VLDB 2026
