FaScalSQL: A Fast and Scalable GPU-Accelerated SQL Query Engine for Out-of-Memory Tables
Chaemin Lim, Suhyun Lee, Jinwoo Choi, Kwanghyun Park, Jinho Lee, Joonsung Kim, Youngsok Kim
摘要
Graphics Processing Units (GPUs) are promising for analytical SQL query processing, but their limited memory capacity hinders processing large input tables exceeding the GPU memory. The existing engines either 1) statically split input columns into chunks and iteratively perform host-to-GPU transfer and relational operations (streaming engines), or 2) maintain a static and fixed-size cache on GPU memory and distribute input columns and their query workloads to the host CPU and a GPU (CPU-GPU distributive engines). However, we find that they suffer from two primary bottlenecks which eventually lead the engines to severe GPU underutilization: excessive host-to-GPU data movement and CPU-GPU load imbalance. We find that they arise from a conflict between their static input data placement and the dynamic progressive filtering of analytical queries. This conflict leads the engines either to transfer column values that are eventually discarded or to assign a large amount of the workload to the host CPU as the input table size scales. In this paper, we present FaScalSQL, a fast and scalable GPU-accelerated SQL query engine that overcomes the severe GPU underutilization of query processing on out-of-memory tables. FaScalSQL introduces a new type of on-demand CPU-GPU coprocessing engine which exploits both GPU-initiated data transfer and CPU-GPU co-processing capability. It replaces the static large unfiltered chunks with a dynamic GPU-initiated on-demand fetching of necessary input data, guided by the host CPU's prefiltering. We evaluate FaScalSQL with Star Schema Benchmark (SSB) and TPC-H. Using SSB with scale factors of 100 and 200, FaScalSQL achieves geometric mean speedups of 2.60× and 2.20× over the existing streaming and CPU-GPU distributive engines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper35
- Cardinality Estimation in DBMS: A Comprehensive Benchmark EvaluationYuxing Han, Ziniu Wu, Peizhi Wu, Rong Zhu 等VLDB 2022 · 被引用 169 次
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 被引用 112 次
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl 等SIGMOD 2020 · 被引用 99 次
- Large Graph Convolutional Network Training with GPU-Oriented Data Communication ArchitectureSeungwon Min, Kun Wu, Sitao Huang, Mert Hidayetoglu 等VLDB 2021 · 被引用 85 次
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu 等VLDB 2021 · 被引用 67 次
相关 Paper
- Scaling GPU-Accelerated Databases beyond GPU Memory SizeYinan Li, Bailu Ding, Ziyun Wei, Lukas M. Maas 等VLDB 2025 · 被引用 7 次
- FineStream: Fine-Grained Window-Based Stream Processing on CPU-GPU Integrated ArchitecturesFeng Zhang, Lin Yang, Shuhao Zhang, Bingsheng He 等USENIX ATC 2020 · 被引用 41 次
- Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and OptimizationKaushik Rajan, Sampath Rajendra, Momin Al-Ghosien, Nicolas Bruno 等VLDB 2026
- Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMSBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2022 · 被引用 45 次
- Terabyte-Scale Analytics in the Blink of an EyeBowen Wu, Wei Cui, Carlo Curino, Matteo Interlandi 等VLDB 2026 · 被引用 10 次
