On-Demand State Separation for Cloud Data Warehousing
Christian Winter, Jana Giceva, Thomas Neumann, Alfons Kemper
Abstract
Moving data analysis and processing to the cloud is no longer reserved for a few companies with petabytes of data. Instead, the flexibility of on-demand resources is attracting an increasing number of customers with small to medium-sized workloads. These workloads do not occupy entire clusters but can run on single worker machines. However, picking the right worker for the job is challenging. Abstracting from worker machines, e.g., using stateless architectures, introduces overheads impacting performance. Solutions without stateless architectures resort to query restarts in the event of an adverse worker matching, wasting already achieved progress. In this paper, we propose migrating queries between workers by introducing on-demand state separation. Using state separation only when required enables maximum flexibility and performance while keeping already achieved progress. To derive the requirements for state separation, we first analyze the query state of medium-sized workloads on the example of TPC-DS SF100. Using this, we analyze the cost and describe the constraints necessary for state separation on such a workload. Furthermore, we describe the design and implementation of on-demand state separation in a compiling database system. Finally, using this implementation, we show the feasibility of our approach on TPC-DS and give a detailed analysis of the cost of query migration and state separation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e5066421-e6fe-477c-9ae1-b73c3c13d3aeCited by top-tier papers4
- Exploiting Cloud Object Storage for High-Performance AnalyticsDominik Durner, Viktor Leis, Thomas NeumannVLDB 2023 · 45 citations
- CXL Memory Performance for In-Memory Data ProcessingMarcel Weisgut, Daniel Ritter, Pinar Tözün, Lawrence Benson et al.VLDB 2025 · 7 citations
- Towards Designing Future-Proof Data Processing SystemsMichael Jungmair, Jana GicevaVLDB 2025 · 1 citation
- Interoperable ACID Transactions for Open Table FormatsTobias Götz, Daniel Ritter, Muhammad El-Hindi, Jana GicevaVLDB 2026
Builds on6
- Building An Elastic Query Engine on Disaggregated StorageMidhul Vuppalapati, Justin Miron, Rachit Agarwal, Dan Truong et al.NSDI 2020 · 142 citations
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu et al.VLDB 2021 · 67 citations
- Towards Cost-Optimal Query Processing in the CloudViktor Leis, Maximilian KuschewskiVLDB 2021 · 34 citations
- Self-Tuning Query Scheduling for Analytical WorkloadsBenjamin Wagner, André Kohn, Thomas NeumannSIGMOD 2021 · 25 citations
- Crystal: A Unified Cache Storage System for Analytical DatabasesDominik Durner, Badrish Chandramouli, Yinan LiVLDB 2021 · 11 citations
Related papers
- Understanding the Effect of Data Center Resource Disaggregation on Production DBMSsQizhen Zhang, Yifan Cai, Xinyi Chen, Sebastian Angel et al.VLDB 2020 · 64 citations
- OPTIMUSCLOUD: Heterogeneous Configuration Optimization for Distributed Databases in the CloudAshraf Mahgoub, Alexander Medoff, Rakesh Kumar, Subrata Mitra et al.USENIX ATC 2020 · 63 citations
- Cloud Analytics BenchmarkAlexander van Renen, Viktor LeisVLDB 2023 · 32 citations
- Rhino: Efficient Management of Very Large Distributed State for Stream Processing EnginesBonaventura Del Monte, Steffen Zeuch, Tilmann Rabl, Volker MarklSIGMOD 2020 · 56 citations
- AQD: Online Adaptive Query Dispatcher for HTAP DatabasesYang Wu, Tongliang Li, Xuanhe Zhou, Jianying Wang et al.VLDB 2026
