NOCAP: Near-Optimal Correlation-Aware Partitioning Joins
Zichen Zhu, Xiao Hu, Manos Athanassoulis
Abstract
Storage-based joins are still commonly used today because the memory budget does not always scale with the data size. One of the many join algorithms developed that has been widely deployed and proven to be efficient is the Hybrid Hash Join (HHJ), which is designed to exploit any available memory to maximize the data that is joined directly in memory. However, HHJ cannot fully exploit detailed knowledge of the join attribute correlation distribution. In this paper, we show that given a correlation skew in the join attributes, HHJ partitions data in a suboptimal way. To do that, we derive the optimal partitioning using a new cost-based analysis of partitioning-based joins that is tailored for primary key -foreign key (PK-FK) joins, one of the most common join types. This optimal partitioning strategy has a high memory cost, thus, we further derive an approximate algorithm that has tunable memory cost and leads to near-optimal results. Our algorithm, termed NOCAP (Near-Optimal Correlation-Aware Partitioning) join, outperforms the state-of-the-art for skewed correlations by up to 30%, and the textbook Grace Hash Join by up to 4×. Further, for a limited memory budget, NOCAP outperforms HHJ by up to 10%, even for uniform correlation. Overall, NOCAP dominates state-of-the-art algorithms and mimics the best algorithm for a memory budget varying from below √︁ ∥relation∥ to more than ∥relation∥. CCS Concepts: • Information systems → Join algorithms; • Theory of computation → Database query processing and optimization (theory).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 79dcbaa7-1a18-4256-8b82-6a3340c7119eCited by top-tier papers2
- High-Performance Query Processing with NVMe Arrays: Spilling without Killing PerformanceMaximilian Kuschewski, Jana Giceva, Thomas Neumann, Viktor LeisSIGMOD 2025 · 11 citations
- Scalable Grid-based Computation of Kendall's Tau CorrelationNikolaos Koutroumanis, Petros Karampas, Alexandros Karakasidis, Nikos Mamoulis et al.VLDB 2026
Builds on5
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2020 · 99 citations
- To Partition, or Not to Partition, That is the Join Question in a Real SystemMaximilian Bandle, Jana Giceva, Thomas NeumannSIGMOD 2021 · 43 citations
- Pushing Data-Induced Predicates Through Joins in Big-Data ClustersLaurel J. Orr, Srikanth Kandula, Surajit ChaudhuriVLDB 2020 · 35 citations
- Cosine: A Cloud-Cost Optimized Self-Designing Key-Value Storage EngineSubarna Chatterjee, Meena Jagadeesan, Wilson Qin, Stratos IdreosVLDB 2022 · 17 citations
- Design Trade-offs for a Robust Dynamic Hybrid Hash JoinShiva Jahangiri, Michael J. Carey, Johann-Christoph FreytagVLDB 2022 · 6 citations
Related papers
- A Design Space Exploration and Evaluation for Main-Memory Hash Joins in Storage Class MemoryWentao Huang, Yunhong Ji, Xuan Zhou, Bingsheng He et al.VLDB 2023 · 9 citations
- Adopting Worst-Case Optimal Joins in Relational Database SystemsMichael J. Freitag, Maximilian Bandle, Tobias Schmidt, Alfons Kemper et al.VLDB 2020 · 79 citations
- Improved Correlated Sampling for Join Size EstimationTaiNing Wang, Chee-Yong ChanICDE 2020 · 19 citations
- FactorJoin: A New Cardinality Estimation Framework for Join QueriesZiniu Wu, Parimarjan Negi, Mohammad Alizadeh, Tim Kraska et al.SIGMOD 2023 · 54 citations
- SOLAR: Scalable Distributed Spatial Joins Through Learning-Based OptimizationYongyi Liu, Ahmed Abdelmaguid, Ahmed R. Mahmood, Amr Magdy et al.ICDE 2026
