Evaluating Multi-GPU Sorting with Modern Interconnects
Tobias Maltenberger, Ivan Ilic, Ilin Tolovski, Tilmann Rabl
摘要
GPUs have become a mainstream accelerator for database operations such as sorting. Most GPU sorting algorithms are single-GPU approaches. They neither harness the full computational power nor exploit the high-bandwidth P2P interconnects of modern multi-GPU platforms. The latest NVLink 2.0 and NVLink 3.0-based NVSwitch interconnects promise unparalleled multi-GPU acceleration. So far, multi-GPU sorting has only been evaluated on systems with PCIe 3.0. In this paper, we analyze serial, parallel, and bidirectional data transfer rates to, from, and between multiple GPUs on systems with PCIe 3.0/4.0, NVLink 2.0/3.0, and NVSwitch. We measure up to 35x higher parallel P2P throughput with NVLink 3.0-based NVSwitch over PCIe 3.0. To study GPU-accelerated sorting on today's hardware, we implement a P2P-based GPU-only (P2P sort) and a heterogeneous (HET sort) multi-GPU sorting algorithm and evaluate them on three modern platforms. We observe speedups over state-of-the-art parallel CPU radix sort of up to 14x for P2P sort and 9x for HET sort. On systems with fast P2P interconnects, P2P sort outperforms HET sort up to 1.65x. Finally, we show that overlapping GPU copy/compute operations does not mitigate the transfer bottleneck when sorting large out-of-core data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- GPU Database Systems Characterization and OptimizationJiashen Cao, Rathijit Sen, Matteo Interlandi, Joy Arulraj 等VLDB 2024 · 被引用 36 次
- Triton Join: Efficiently Scaling to a Large Join State on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl 等SIGMOD 2022 · 被引用 24 次
- BOSS - An Architecture for Database Kernel CompositionHubert Mohr-Daurat, Xuan Sun, Holger PirkVLDB 2024 · 被引用 12 次
- Analyzing Vectorized Hash Tables Across CPU ArchitecturesMaximilian Böther, Lawrence Benson, Ana Klimovic, Tilmann RablVLDB 2023 · 被引用 11 次
- Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data AnalyticsYichao Yuan, Advait Iyer, Lin Ma, Nishil TalatiVLDB 2025 · 被引用 11 次
它引用的顶会 Paper3
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl 等SIGMOD 2020 · 被引用 99 次
- Efficient Join Algorithms For Large Database Tables in a Multi-GPU EnvironmentRan Rui, Hao Li, Yi-Cheng TuVLDB 2021 · 被引用 43 次
- MG-Join: A Scalable Join for Massively Parallel Multi-GPU ArchitecturesJohns Paul, Shengliang Lu, Bingsheng He, Chiew Tong LauSIGMOD 2021 · 被引用 31 次
相关 Paper
- Efficiently Joining Large Relations on Multi-GPU SystemsTobias Maltenberger, Ilin Tolovski, Tilmann RablVLDB 2025 · 被引用 3 次
- Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU SystemsChen Zhang, Qijun Zhang, Zhuoshan Zhou, Yijia Diao 等HPCA 2026 · 被引用 1 次
- Distributed GPU Joins on Fast RDMA-capable NetworksLasse Thostrup, Gloria Doci, Nils Boeschen, Manisha Luthra 等SIGMOD 2023 · 被引用 18 次
- Poplar: Efficient Scaling of Distributed DNN Training on Heterogeneous GPU ClustersWenZheng Zhang, Yang Hu, Jing Shi, Xiaoying BaiAAAI 2025 · 被引用 5 次
- PPipe: Efficient Video Analytics Serving on Heterogeneous GPU Clusters via Pool-Based Pipeline ParallelismZ. Jonny Kong, Qiang Xu, Y. Charlie HuUSENIX ATC 2025 · 被引用 4 次
