Scalable Multi-Query Execution using Reinforcement Learning
Panagiotis Sioulas, Anastasia Ailamaki
Abstract
The growing demand for data-intensive decision support and the migration to multi-tenant infrastructures put databases under the stress of high analytical query load. The requirement for high throughput contradicts the traditional design of query-at-a-time databases that optimize queries for efficient serial execution. Sharing work across queries presents an opportunity to reduce the total cost of processing and therefore improve throughput with increasing query load. Systems can share work either by assessing all opportunities and restructuring batches of queries ahead of execution, or by inspecting opportunities in individual incoming queries at runtime: the former strategy scales poorly to large query counts, as it requires expensive sharing-aware optimization, whereas the latter detects only a subset of the opportunities. Both strategies fail to minimize the cost of processing for large and ad-hoc workloads. This paper presents RouLette, a specialized intelligent engine for multi-query execution that addresses, through runtime adaptation, the shortcomings of existing work-sharing strategies. RouLette scales by replacing sharing-aware optimization with adaptive query processing, and it chooses opportunities to explore and exploit by using reinforcement learning. RouLette also includes optimizations that reduce the adaptation overhead. RouLette increases throughput by 1.6-28.3x, compared to a state-of-the-art query-at-a-time engine, and up to 6.5x, compared to sharing-enabled prototypes, for multi-query workloads based on the schema of TPC-DS.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 375b0633-2537-4b57-afaa-04628f94071eCited by top-tier papers3
- LSched: A Workload-Aware Learned Query Scheduler for Analytical Database SystemsIbrahim Sabek, Tenzin Samten Ukyab, Tim KraskaSIGMOD 2022 · 25 citations
- Using Cloud Functions as Accelerator for Elastic Data AnalyticsHaoqiong Bian, Tiannan Sha, Anastasia AilamakiSIGMOD 2023 · 18 citations
- Process Faster, Pay Less: Functional Isolation for Stream ProcessingEleni Zapridou, Michael Koepf, Panagiotis Sioulas, Ioannis Mytilinis et al.ICDE 2026
Related papers
- Simple Adaptive Query Processing vs. Learned Query Optimizers: Observations and AnalysisYunjia Zhang, Yannis Chronis, Jignesh M. Patel, Theodoros RekatsinasVLDB 2023 · 20 citations
- ADOPT: Adaptively Optimizing Attribute Orders for Worst-Case Optimal Join Algorithms via Reinforcement LearningJunxiong Wang, Immanuel Trummer, Ahmet Kara, Dan OlteanuVLDB 2023 · 10 citations
- SASPAR: Shared Adaptive Stream PartitioningJeyhun Karimov, Hans-Arno JacobsenICDE 2023 · 3 citations
- SkinnerMT: Parallelizing for Efficiency and Robustness in Adaptive Query Processing on Multicore PlatformsZiyun Wei, Immanuel TrummerVLDB 2023 · 4 citations
- QaaD (Query-as-a-Data): Scalable Execution of Massive Number of Small Queries in SparkYeonsu Park, Byungchul Tak, Wook-Shin HanSIGMOD 2023 · 3 citations
