MG-Join: A Scalable Join for Massively Parallel Multi-GPU Architectures
Johns Paul, Shengliang Lu, Bingsheng He, Chiew Tong Lau
Abstract
The recent scale-up of GPU hardware through the integration of multiple GPUs into a single machine and the introduction of higher bandwidth interconnects like NVLink 2.0 have enabled new opportunities of relational query processing on multiple GPUs. However, due to the unique characteristics of GPUs and the interconnects, existing hash join implementations spend up to 66% of their execution time moving the data between the GPUs and achieve lower than 50% utilization of the newer high bandwidth interconnects. This leads to extremely poor scalability of hash join performance on multiple GPUs, which can be slower than the performance on a single GPU. In this paper, we propose MG-Join, a scalable partitioned hash join implementation on multiple GPUs of a single machine. In order to effectively improve the bandwidth utilization, we develop a novel multi-hop routing for cross-GPU communication that adaptively chooses the efficient route for each data flow to minimize congestion. Our experiments on the DGX-1 machine show that MG-Join helps significantly reduce the communication overhead and achieves up to 97% utilization of the bisection bandwidth of the interconnects, resulting in significantly better scalability. Overall, MG-Join outperforms the state-of-the-art hash join implementations by up to 2.5x. MG-Join further helps improve the overall performance of TPC-H queries by up to 4.5x over multi-GPU version of an open-source commercial GPU database Omnisci.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Tile-based Lightweight Integer Compression in GPUAnil Shanbhag, Bobbi W. Yogatama, Xiangyao Yu, Samuel MaddenSIGMOD 2022 · 45 citations
- Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMSBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2022 · 45 citations
- GPU Database Systems Characterization and OptimizationJiashen Cao, Rathijit Sen, Matteo Interlandi, Joy Arulraj et al.VLDB 2024 · 36 citations
- RTIndeX: Exploiting Hardware-Accelerated GPU Raytracing for Database IndexingJustus Henneberg, Felix SchuhknechtVLDB 2023 · 25 citations
- Evaluating Multi-GPU Sorting with Modern InterconnectsTobias Maltenberger, Ivan Ilic, Ilin Tolovski, Tilmann RablSIGMOD 2022 · 24 citations
Builds on2
Related papers
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2020 · 99 citations
- Efficiently Joining Large Relations on Multi-GPU SystemsTobias Maltenberger, Ilin Tolovski, Tilmann RablVLDB 2025 · 3 citations
- Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2022 · 24 citations
- Efficient Join Algorithms For Large Database Tables in a Multi-GPU EnvironmentRan Rui, Hao Li, Yi-Cheng TuVLDB 2021 · 43 citations
- Distributed GPU Joins on Fast RDMA-capable NetworksLasse Thostrup, Gloria Doci, Nils Boeschen, Manisha Luthra et al.SIGMOD 2023 · 18 citations
