FaScalSQL: A Fast and Scalable GPU-Accelerated SQL Query Engine for Out-of-Memory Tables
Chaemin Lim, Suhyun Lee, Jinwoo Choi, Kwanghyun Park, Jinho Lee, Joonsung Kim, Youngsok Kim
Abstract
Graphics Processing Units (GPUs) are promising for analytical SQL query processing, but their limited memory capacity hinders processing large input tables exceeding the GPU memory. The existing engines either 1) statically split input columns into chunks and iteratively perform host-to-GPU transfer and relational operations (streaming engines), or 2) maintain a static and fixed-size cache on GPU memory and distribute input columns and their query workloads to the host CPU and a GPU (CPU-GPU distributive engines). However, we find that they suffer from two primary bottlenecks which eventually lead the engines to severe GPU underutilization: excessive host-to-GPU data movement and CPU-GPU load imbalance. We find that they arise from a conflict between their static input data placement and the dynamic progressive filtering of analytical queries. This conflict leads the engines either to transfer column values that are eventually discarded or to assign a large amount of the workload to the host CPU as the input table size scales. In this paper, we present FaScalSQL, a fast and scalable GPU-accelerated SQL query engine that overcomes the severe GPU underutilization of query processing on out-of-memory tables. FaScalSQL introduces a new type of on-demand CPU-GPU coprocessing engine which exploits both GPU-initiated data transfer and CPU-GPU co-processing capability. It replaces the static large unfiltered chunks with a dynamic GPU-initiated on-demand fetching of necessary input data, guided by the host CPU's prefiltering. We evaluate FaScalSQL with Star Schema Benchmark (SSB) and TPC-H. Using SSB with scale factors of 100 and 200, FaScalSQL achieves geometric mean speedups of 2.60× and 2.20× over the existing streaming and CPU-GPU distributive engines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3386dca2-8261-48d3-9ae0-38ae47b6e181Builds on35
- Cardinality Estimation in DBMS: A Comprehensive Benchmark EvaluationYuxing Han, Ziniu Wu, Peizhi Wu, Rong Zhu et al.VLDB 2022 · 169 citations
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 112 citations
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2020 · 99 citations
- Large Graph Convolutional Network Training with GPU-Oriented Data Communication ArchitectureSeungwon Min, Kun Wu, Sitao Huang, Mert Hidayetoglu et al.VLDB 2021 · 85 citations
- FlexPushdownDB: Hybrid Pushdown and Caching in a Cloud DBMSYifei Yang, Matt Youill, Matthew E. Woicik, Yizhou Liu et al.VLDB 2021 · 67 citations
Related papers
- Scaling GPU-Accelerated Databases beyond GPU Memory SizeYinan Li, Bailu Ding, Ziyun Wei, Lukas M. Maas et al.VLDB 2025 · 7 citations
- FineStream: Fine-Grained Window-Based Stream Processing on CPU-GPU Integrated ArchitecturesFeng Zhang, Lin Yang, Shuhao Zhang, Bingsheng He et al.USENIX ATC 2020 · 41 citations
- Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and OptimizationKaushik Rajan, Sampath Rajendra, Momin Al-Ghosien, Nicolas Bruno et al.VLDB 2026
- Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMSBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2022 · 45 citations
- Terabyte-Scale Analytics in the Blink of an EyeBowen Wu, Wei Cui, Carlo Curino, Matteo Interlandi et al.VLDB 2026 · 10 citations
