Efficiently Joining Large Relations on Multi-GPU Systems
Tobias Maltenberger, Ilin Tolovski, Tilmann Rabl
Abstract
Growing data volumes present a mounting challenge to relational joins. GPUs have gained widespread adoption as database accelerators for operators such as joins due to their high instruction throughput and memory bandwidth. Most published GPU-accelerated joins are single-GPU algorithms that do not leverage modern multi-GPU platforms effectively. The few proposed multi-GPU algorithms either fail to exploit the high-speed P2P interconnects between the GPUs or to handle large out-of-core data natively. In this paper, we present a heterogeneous multi-GPU sort-merge join that overcomes both limitations. It is composed of a merge- or radix partitioning-based P2P-enabled multi-GPU sort phase, a parallel CPU-based multiway merge phase, and a hybrid join phase that combines a CPU merge path partition with a binary search-based multi-GPU join strategy. We evaluate our novel multi-GPU join on two platforms with fast NVLink- and NVSwitch-based P2P interconnects. We show that our join outperforms state-of-the-art CPU and GPU baselines regardless of the workload. It outperforms parallel CPU sort-merge and radix-hash joins by up to 15.2× and 5.5×, respectively. Compared to non-P2P-enabled multi-GPU joins, it achieves speedups of 8.7× (sort-merge) and 2.5× (hybrid-radix). We measure that our join's hybrid join phase with overlapped copy and compute operations contributes as little as 22% to its end-to-end runtime. If the input relations are pre-sorted, it is up to 14.4× faster than the hybrid-radix join. Our join scales well with the number of GPUs and benefits from data skew with as much as 12% shorter join durations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 707aa772-4f5a-4af1-ad31-947f1446cbfbBuilds on11
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 112 citations
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2020 · 99 citations
- Efficient Join Algorithms For Large Database Tables in a Multi-GPU EnvironmentRan Rui, Hao Li, Yi-Cheng TuVLDB 2021 · 43 citations
- To Partition, or Not to Partition, That is the Join Question in a Real SystemMaximilian Bandle, Jana Giceva, Thomas NeumannSIGMOD 2021 · 43 citations
- MG-Join: A Scalable Join for Massively Parallel Multi-GPU ArchitecturesJohns Paul, Shengliang Lu, Bingsheng He, Chiew Tong LauSIGMOD 2021 · 31 citations
Related papers
- Evaluating Multi-GPU Sorting with Modern InterconnectsTobias Maltenberger, Ivan Ilic, Ilin Tolovski, Tilmann RablSIGMOD 2022 · 24 citations
- Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2022 · 24 citations
- Distributed GPU Joins on Fast RDMA-capable NetworksLasse Thostrup, Gloria Doci, Nils Boeschen, Manisha Luthra et al.SIGMOD 2023 · 18 citations
- Efficiently Processing Joins and Grouped Aggregations on GPUsBowen Wu, Dimitrios Koutsoukos, Gustavo AlonsoSIGMOD 2025 · 15 citations
- A Case for Graphics-driven Query ProcessingHarish Doraiswamy, Vikas Kalagi, Karthik Ramachandra, Jayant R. HaritsaVLDB 2023 · 4 citations
