Incorporating Super-Operators in Big-Data Query Optimizers
Jyoti Leeka, Kaushik Rajan
摘要
The cost of big-data analytics is dominated by shuffle operations that induce multiple disk reads, writes and network transfers. This paper proposes a new class of optimization rules that are specifically aimed at eliminating shuffles where possible. The rules substitute multiple shuffle inducing operators ( Join, UnionAll, Spool, GroupBy ) with a single streaming operator which implements an entire sub-query. We call such operators super-operators. A key challenge with adding new rules that substitute sub-queries with super-operators is that there are many variants of the same sub-query that can be implemented via minor modifications to the same super-operator. Adding each as a separate rule leads to a search space explosion. We propose several extensions to the query optimizer to address this challenge. We propose a new abstract representation for operator trees that captures all possible sub-queries that a super-operator implements. We propose a new rule matching algorithm that can efficiently search for abstract operator trees. Finally we extend the physical operator interface to introduce new parametric super-operators. We implement our changes in SCOPE, a state-of-the-art production big-data optimizer used extensively at Microsoft. We demonstrate that the proposed optimizations provide significant reduction in both resource cost (average 1.7x) and latency (average 1.5x) on several production queries, and do so without increasing optimization time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- GenRewrite: Query Rewriting via Large Language ModelsJie Liu, Barzan MozafariSIGMOD 2026 · 被引用 27 次
- Predicate Pushdown for Data Science PipelinesCong Yan, Yin Lin, Yeye HeSIGMOD 2023 · 被引用 15 次
- SlabCity: Whole-Query Optimization using Program SynthesisRui Dong, Jie Liu, Yuxuan Zhu, Cong Yan 等VLDB 2023 · 被引用 11 次
- Phoebe: A Learning-based Checkpoint OptimizerYiwen Zhu, Matteo Interlandi, Abhishek Roy, Krishnadhan Das 等VLDB 2021 · 被引用 10 次
- Modularis: Modular Relational Analytics over Heterogeneous Distributed PlatformsDimitrios Koutsoukos, Ingo Müller, Renato Marroquín, Ana Klimovic 等VLDB 2021 · 被引用 8 次
相关 Paper
- New Query Optimization Techniques in the Spark Engine of Azure SynapseAbhishek Modi, Kaushik Rajan, Srinivas Thimmaiah, Prakhar Jain 等VLDB 2022 · 被引用 13 次
- Generalized Sub-Query Fusion for Eliminating Redundant I/O from Big-Data QueriesPartho Sarthi, Kaushik Rajan, Akash Lal, Abhishek Modi 等OSDI 2020 · 被引用 5 次
- Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsTarique Siddiqui, Alekh Jindal, Shi Qiao, Hiren Patel 等SIGMOD 2020 · 被引用 80 次
- SASPAR: Shared Adaptive Stream PartitioningJeyhun Karimov, Hans-Arno JacobsenICDE 2023 · 被引用 3 次
- COMPARE: Accelerating Groupwise Comparison in Relational Databases for Data AnalyticsTarique Siddiqui, Surajit Chaudhuri, Vivek R. NarasayyaVLDB 2021 · 被引用 19 次
