To Partition, or Not to Partition, That is the Join Question in a Real System
Maximilian Bandle, Jana Giceva, Thomas Neumann
摘要
An efficient implementation of a hash join has been a highly researched problem for decades. Recently, the radix join has been shown to have superior performance over the alternatives (e.g., the non-partitioned hash join), albeit on synthetic microbenchmarks. Therefore, it is unclear whether one can simply replace the hash join in an RDBMS or use the radix join as a performance booster for selected queries. If the latter, it is still unknown when one should rely on the radix join to improve performance.
In this paper, we address these questions, show how to integrate the radix join in Umbra, a code-generating DBMS, and make it competitive for selective queries by introducing a Bloom-filter based semi-join reducer. We have evaluated how well it runs when used in queries from more representative workloads like TPC-H. Surprisingly, the radix join brings a noticeable improvement in only one out of all 59 joins in TPC-H. Thus, with an extensive range of microbenchmarks, we have isolated the effects of the most important workload factors and synthesized the range of values where partitioning the data for the radix join pays off. Our analysis shows that the benefit of data partitioning quickly diminishes as soon as we deviate from the optimal parameters, and even late materialization rarely helps in real workloads. We thus, conclude that integrating the radix join within a code-generating database rarely justifies the increase in code and optimizer complexity and advise against it for processing real-world workloads.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- The Case for Learned In-Memory JoinsIbrahim Sabek, Tim KraskaVLDB 2023 · 被引用 27 次
- Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl 等SIGMOD 2022 · 被引用 24 次
- Fast Detection of Denial Constraint ViolationsEduardo H. M. Pena, Eduardo Cunha de Almeida, Felix NaumannVLDB 2022 · 被引用 22 次
- Database Technology for the Masses: Sub-Operators as First-Class EntitiesMaximilian Bandle, Jana GicevaVLDB 2021 · 被引用 19 次
- Efficiently Processing Joins and Grouped Aggregations on GPUsBowen Wu, Dimitrios Koutsoukos, Gustavo AlonsoSIGMOD 2025 · 被引用 15 次
它引用的顶会 Paper1
相关 Paper
- A Scalable and Generic Approach to Range JoinsMaximilian Reif, Thomas NeumannVLDB 2022 · 被引用 6 次
- Design Trade-offs for a Robust Dynamic Hybrid Hash JoinShiva Jahangiri, Michael J. Carey, Johann-Christoph FreytagVLDB 2022 · 被引用 6 次
- Bringing Compiling Databases to RISC ArchitecturesFerdinand Gruber, Maximilian Bandle, Alexis Engelke, Thomas Neumann 等VLDB 2023 · 被引用 9 次
- Robust Predicate Transfer with Dynamic ExecutionYiming Qiao, Peter Boncz, Huanchen ZhangVLDB 2026 · 被引用 2 次
- Detecting Join Bugs in Database Engines via Join Implication ReasoningZhaokun Xiang, Suyang Zhong, Manuel RiggerSIGMOD 2026
