Lune

VLDB2026Top-tier venue

Terabyte-Scale Analytics in the Blink of an Eye

Bowen Wu, Wei Cui, Carlo Curino, Matteo Interlandi, Rathijit Sen

2026Year
10Citations
5Top-tier citations

Abstract

For the past two decades, the DB community has devoted substantial research to take advantage of cheap clusters of machines for distributed data analytics-we believe that we are at the beginning of a paradigm shift. The scaling laws and popularity of AI models lead to the deployment of incredibly powerful GPU clusters in commercial data centers. Compared to CPU-only solutions, these clusters deliver impressive improvements in per-node compute, memory bandwidth, and inter-node interconnect performance. In this paper, we study the problem of scaling analytical SQL queries on distributed clusters of GPUs, with the stated goal of establishing an upper bound on the likely performance gains. To do so, we build a prototype designed to maximize performance by leveraging ML/HPC best practices, such as group communication primitives for cross-device data movements. This allows us to conduct thorough performance experimentation to point our community towards a massive performance opportunity of at least 60×. To make these gains more relatable, before you can blink twice, our system can run all 22 queries of TPC-H at a 1TB scale factor! * Work done during an internship at the Microsoft Gray Systems Lab. † Equal contribution.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ebb2a329-835e-4544-91ff-cd24fb42ff3c

Cited by top-tier papers5

Ask how each one uses it

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines