Terabyte-Scale Analytics in the Blink of an Eye
Bowen Wu, Wei Cui, Carlo Curino, Matteo Interlandi, Rathijit Sen
摘要
For the past two decades, the DB community has devoted substantial research to take advantage of cheap clusters of machines for distributed data analytics-we believe that we are at the beginning of a paradigm shift. The scaling laws and popularity of AI models lead to the deployment of incredibly powerful GPU clusters in commercial data centers. Compared to CPU-only solutions, these clusters deliver impressive improvements in per-node compute, memory bandwidth, and inter-node interconnect performance. In this paper, we study the problem of scaling analytical SQL queries on distributed clusters of GPUs, with the stated goal of establishing an upper bound on the likely performance gains. To do so, we build a prototype designed to maximize performance by leveraging ML/HPC best practices, such as group communication primitives for cross-device data movements. This allows us to conduct thorough performance experimentation to point our community towards a massive performance opportunity of at least 60×. To make these gains more relatable, before you can blink twice, our system can run all 22 queries of TPC-H at a 1TB scale factor! * Work done during an internship at the Microsoft Gray Systems Lab. † Equal contribution.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- High-Performance DBMSs with io_uring: When and How to Use ItMatthias Jasny, Muhammad El-Hindi, Tobias Ziegler, Viktor Leis 等VLDB 2026 · 被引用 5 次
- PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast StorageJigao Luo, Nils Boeschen, Muhammad El-Hindi, Carsten BinnigVLDB 2026
- FaScalSQL: A Fast and Scalable GPU-Accelerated SQL Query Engine for Out-of-Memory TablesChaemin Lim, Suhyun Lee, Jinwoo Choi, Kwanghyun Park 等ICDE 2026
- MGI: A Communication Framework for Data Processing in Massive GPU InfrastructuresDi Wu, Hongshi Tan, Hanzhang Yang, Bingsheng He 等VLDB 2026
- Understanding and Profiling the Accelerator Chiplet Network Using PingPointJunyeol Ryu, Ming Liu, Matthew D. SinclairSIGCOMM 2026
它引用的顶会 Paper27
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen 等ICLR 2021 · 被引用 1,954 次
- TopoOpt: Co-optimizing Network Topology and Parallelization Strategy for Distributed Training JobsWeiyang Wang, Moein Khazraee, Zhizhen Zhong, Manya Ghobadi 等NSDI 2023 · 被引用 215 次
- Efficient sparse collective communication and its application to accelerate distributed deep learningJiawei Fei, Chen-Yu Ho, Atal Narayan Sahu, Marco Canini 等SIGCOMM 2021 · 被引用 120 次
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 被引用 112 次
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl 等SIGMOD 2020 · 被引用 99 次
相关 Paper
- Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and OptimizationKaushik Rajan, Sampath Rajendra, Momin Al-Ghosien, Nicolas Bruno 等VLDB 2026
- GPU Database Systems Characterization and OptimizationJiashen Cao, Rathijit Sen, Matteo Interlandi, Joy Arulraj 等VLDB 2024 · 被引用 36 次
- Scaling GPU-Accelerated Databases beyond GPU Memory SizeYinan Li, Bailu Ding, Ziyun Wei, Lukas M. Maas 等VLDB 2025 · 被引用 7 次
- Scaling your Hybrid CPU-GPU DBMS to Multiple GPUsBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2024 · 被引用 9 次
- Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMSBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2022 · 被引用 45 次
