PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast Storage
Jigao Luo, Nils Boeschen, Muhammad El-Hindi, Carsten Binnig
Abstract
The AI hardware boom has led modern data centers to adopt HPCstyle architectures centered on distributed, GPU-centric computation. Large GPU clusters interconnected by fast RDMA networks and backed by high-bandwidth NVMe storage enable scalable computation and rapid access to storage-resident data. Tensor computation runtimes (TCRs), such as PyTorch, originally designed for AI workloads, have recently been shown to accelerate analytical workloads. However, prior work has primarily considered settings where the data ts in aggregated GPU memory. In this paper, we systematically study how TCRs can support scalable, distributed query processing for large-scale, storage-resident OLAP workloads. Although TCRs provide abstractions for network and storage I/O, naive use often underutilizes GPU and I/O bandwidth due to insufcient overlap between computation and data movement. As a core contribution, we present PystachIO, a prototype of a PyTorch-based distributed OLAP engine that combines fast network and storage I/O with key optimizations to maximize GPU, network, and storage utilization. Our evaluation shows up to 3⇥ end-to-end speedups over existing distributed GPU-based query processing approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2efd0837-c015-4cbe-aef4-e16253937013Builds on13
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 112 citations
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2020 · 99 citations
- Query Processing on Tensor Computation RuntimesDong He, Supun Chathuranga Nakandala, Dalitso Banda, Rathijit Sen et al.VLDB 2022 · 54 citations
- ScaleStore: A Fast and Cost-Efficient Storage Engine using DRAM, NVMe, and RDMATobias Ziegler, Carsten Binnig, Viktor LeisSIGMOD 2022 · 51 citations
- BtrBlocks: Efficient Columnar Compression for Data LakesMaximilian Kuschewski, David Sauerwein, Adnan Alhomssi, Viktor LeisSIGMOD 2023 · 47 citations
Related papers
- TenGraph: A Tensor-Based Graph Query EngineGuanghua Li, Hao Zhang, Xibo Sun, Qiong Luo et al.VLDB 2024 · 4 citations
- TQEx: Tensor-based Query Engine Enhanced by Bridging the GapHaitao Zhang, Ran Pang, Yuanyuan Zhu, Hao Zhang et al.SIGMOD 2026
- Tensor Relational Algebra for Distributed Machine Learning System DesignBinhang Yuan, Dimitrije Jankov, Jia Zou, Yuxin Tang et al.VLDB 2021 · 33 citations
- Improving Execution Efficiency of Just-in-time Compilation based Query Processing on GPUsJohns Paul, Bingsheng He, Shengliang Lu, Chiew Tong LauVLDB 2021 · 28 citations
- CoPilotIO: CPU as a Co-Pilot for GPU I/O to Free GPU ComputeGuanyi Chen, Qi Chen, Shu Yin, Jian ZhangOSDI 2026
