Improving Execution Efficiency of Just-in-time Compilation based Query Processing on GPUs
Johns Paul, Bingsheng He, Shengliang Lu, Chiew Tong Lau
Abstract
In recent years, we have witnessed significant efforts to improve the performance of Online Analytical Processing (OLAP) on graphics processing units (GPUs). Most existing studies have focused on improving memory efficiency since memory stalls can play an essential role in query processing performance on GPUs. Motivated by the recent rise of just-in-time (JIT) compilation in query processing, we investigate whether and how we can further improve query processing performance on GPU. Specifically, we study the execution of state-of-the-art JIT compile-based query processing systems. We find that thanks to advanced techniques such as database compression and JIT compilation, memory stalls are no longer the most significant bottleneck. Instead, current JIT compile-based query processing encounters severe under-utilization of GPU hardware due to divergent execution and degraded parallelism arising from resource contention. To address these issues, we propose a JIT compile-based query engine named Pyper to improve GPU utilization during query execution. Specifically, Pyper has two new operators, Shuffle and Segment , for query plan transformation, which can be plugged into a physical query plan in order to reduce divergent execution and resolve resource contention, respectively. To determine the insertion points for these two operators, we present an analytical model that helps insert Shuffle and Segment operators into a query plan in a cost-based manner. Our experiments show that 1) the analytical analysis of divergent execution and resource contention helps to improve the accuracy of the cost model, 2) Pyper significantly outperforms other GPU query engines on TPC-H and SSB queries.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8125364-1945-4aaa-b8f4-e8f4c514968bCited by top-tier papers10
- Query Processing on Tensor Computation RuntimesDong He, Supun Chathuranga Nakandala, Dalitso Banda, Rathijit Sen et al.VLDB 2022 · 54 citations
- GPU Database Systems Characterization and OptimizationJiashen Cao, Rathijit Sen, Matteo Interlandi, Joy Arulraj et al.VLDB 2024 · 36 citations
- MG-Join: A Scalable Join for Massively Parallel Multi-GPU ArchitecturesJohns Paul, Shengliang Lu, Bingsheng He, Chiew Tong LauSIGMOD 2021 · 31 citations
- BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler ApproachZhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu et al.SIGMOD 2024 · 14 citations
- Scaling your Hybrid CPU-GPU DBMS to Multiple GPUsBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2024 · 9 citations
Builds on4
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 112 citations
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2020 · 99 citations
- Data-Parallel Query Processing on Non-Uniform DataHenning Funke, Jens TeubnerVLDB 2020 · 34 citations
- Getting Swole: Generating Access-Aware Code with Predicate PullupsAndrew Crotty, Alex Galakatos, Tim KraskaICDE 2020 · 11 citations
Related papers
- Permutable Compiled Queries: Dynamically Adapting Compiled Queries without RecompilingPrashanth Menon, Amadou Ngom, Todd C. Mowry, Andrew Pavlo et al.VLDB 2021 · 26 citations
- PystachIO: Efficient Distributed GPU Query Processing with PyTorch over Fast Networks & Fast StorageJigao Luo, Nils Boeschen, Muhammad El-Hindi, Carsten BinnigVLDB 2026
- A Case for Graphics-driven Query ProcessingHarish Doraiswamy, Vikas Kalagi, Karthik Ramachandra, Jayant R. HaritsaVLDB 2023 · 4 citations
- GPU Acceleration of SQL Analytics on Compressed DataZezhou Huang, Krystian Sakowski, Hans Lehnert, Wei Cui et al.VLDB 2026 · 1 citation
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 13 citations
