Salvaging failing and straggling queries
Bruhathi Sundarmurthy, Harshad Deshmukh, Paris Koutris, Jeffrey F. Naughton
Abstract
Interactive time responses are a crucial requirement for users analyzing large amounts of data, typically stored in a relational style data-warehouse where data is partitioned across thousands of nodes for high efficiency and throughput. However, consistently providing quick responses remains a big challenge for two reasons: (1) with data distributed across thousands of nodes, it is highly likely that some nodes are unavailable or are very slow during query execution and, (2) large number of users result in high resource contention which exacerbates the problem of slow and failing nodes. In such situations, systems typically straggle or fail the query resulting in higher latencies and wastage of resources. In this paper, we propose a novel solution to alleviate the failure/straggling problem: use the intermediate results from the partial query execution over available data, and exploit the statistical properties of efficiently partitioned data, particularly, co-hash partitioned data, to provide approximate answers along with confidence bounds. The proposed approach handles aggregate queries that involve joins, group bys, having clauses and a subclass of nested subqueries, covering a large portion of analytical queries. We validate our approach through extensive experiments on the TPC-H dataset and we observe that even with a low data availability of 1%, our proposed solution provides answers with less than 5% error.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 3ac189d9-d326-4015-bf94-b650d64e5755Related papers
- Approximate Partition Selection for Big-Data Workloads using Summary StatisticsKexin Rong, Yao Lu, Peter Bailis, Srikanth Kandula et al.VLDB 2020 · 8 citations
- Rotary: A Resource Arbitration Framework for Progressive Iterative AnalyticsRui Liu, Aaron J. Elmore, Michael J. Franklin, Sanjay KrishnanICDE 2023 · 1 citation
- Rapid Approximate Aggregation with Distribution-Sensitive Interval GuaranteesStephen Macke, Maryam Aliakbarpour, Ilias Diakonikolas, Aditya G. Parameswaran et al.ICDE 2021 · 3 citations
- PAW: Data Partitioning Meets Workload VarianceZhe Li, Man Lung Yiu, Tsz Nam ChanICDE 2022 · 6 citations
- Conditional Generative Model Based Predicate-Aware Query ApproximationNikhil Sheoran, Subrata Mitra, Vibhor Porwal, Siddharth Ghetia et al.AAAI 2022 · 14 citations
