Scalable GPU Acceleration of Scalar Functions in Analytical Databases: Compilation, Benchmarking, and Optimization
Kaushik Rajan, Sampath Rajendra, Momin Al-Ghosien, Nicolas Bruno, Carlo Curino, Matteo Interlandi, Yinan Li, Lukas M. Maas, Craig Peeper, Surajit Chaudhuri, Johannes Gehrke
摘要
Accelerating SQL query execution with GPUs is a central focus in database research. While prior systems have achieved notable speedups by offloading relational operators, the acceleration of the wide range of scalar functions that are supported by analytical engines remains unaddressed. Our analysis reveals that many scalar functions incur substantial computational overhead and often constitute the primary bottleneck in analytical queries on CPUs. This observation motivates a systematic exploration of the opportunities and challenges in accelerating scalar functions on GPUs.
Unlike relational operators, which are few in number and standardized, production databases support hundreds of scalar functions. The absence of a standardized specification, combined with this diversity, renders manual GPU porting infeasible. To address this, we present an LLVM-MLIR-based compiler toolchain that automatically translates the CPU-based implementations of scalar functions from production databases into efficient GPU kernels, while preserving their original semantics. Our approach lifts scalar functions to a high-level intermediate representation, applies resource-optimizing transformations, and generates GPU assembly code, supporting all relevant data types, parameters, and database context variables.
As existing benchmarks do not sufficiently stress test scalar functions in analytical queries, we introduce a variant of TPC-H that utilizes scalar functions while preserving the original query intent. Integrating our GPU kernels into a state-of-the-art GPU database system, we demonstrate substantial performance gains over a leading CPU database that uses slightly more expensive hardware: 7.6× on enhanced TPC-H and 6.4× on production queries, further widening the gap between GPU and CPU databases. The generated kernels deliver performance comparable to hand-optimized GPU implementations, establishing our approach as a scalable and practical solution for accelerating scalar functions on GPUs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 被引用 112 次
- Query Processing on Tensor Computation RuntimesDong He, Supun Chathuranga Nakandala, Dalitso Banda, Rathijit Sen 等VLDB 2022 · 被引用 54 次
- Designing an Open Framework for Query Optimization and CompilationMichael Jungmair, André Kohn, Jana GicevaVLDB 2022 · 被引用 45 次
- Efficient Join Algorithms For Large Database Tables in a Multi-GPU EnvironmentRan Rui, Hao Li, Yi-Cheng TuVLDB 2021 · 被引用 43 次
- Tensors: An abstraction for general data processingDimitrios Koutsoukos, Supun Nakandala, Konstantinos Karanasos, Karla Saur 等VLDB 2021 · 被引用 38 次
相关 Paper
- FaScalSQL: A Fast and Scalable GPU-Accelerated SQL Query Engine for Out-of-Memory TablesChaemin Lim, Suhyun Lee, Jinwoo Choi, Kwanghyun Park 等ICDE 2026
- A Case for Graphics-driven Query ProcessingHarish Doraiswamy, Vikas Kalagi, Karthik Ramachandra, Jayant R. HaritsaVLDB 2023 · 被引用 4 次
- Terabyte-Scale Analytics in the Blink of an EyeBowen Wu, Wei Cui, Carlo Curino, Matteo Interlandi 等VLDB 2026 · 被引用 10 次
- GPU Database Systems Characterization and OptimizationJiashen Cao, Rathijit Sen, Matteo Interlandi, Joy Arulraj 等VLDB 2024 · 被引用 36 次
- Data-Parallel Query Processing on Non-Uniform DataHenning Funke, Jens TeubnerVLDB 2020 · 被引用 34 次
