Bridging the Indexing Gap in Fused GPU Query Engines
Tianjun Bu, Gaoyuan Zhou, Xuhui Li, Qiusong Yang
Abstract
GPU query backends achieve high throughput on analytical workloads through massive parallelism, but lack indexing support that accelerates selective queries in CPU databases. Existing GPU index implementations face three limitations: (1) supporting conjunctive predicates only, (2) materializing intermediate results between index access and query execution, and (3) assuming query boundaries align with pre-built index bins. We present a fused bitmap indexing approach that addresses these limitations. We introduce virtual query program that executes arbitrary boolean predicates with low overhead. We fuse index access with subsequent column lookups, joins, and aggregation, keeping intermediate results in registers and eliminating global-memory round-trips. To handle misaligned query boundaries, we propose GPU friendly candidate checking that tracks three-valued row states (certain-in, certain-out, uncertain) through in-register boolean operations and verifies only the necessary candidates, without accessing global memory. On Star Schema Benchmark SF=140 with RTX 5090 D, our fused bitmap index achieves up to 6.9× geometric-mean speed over our optimized non-indexed baseline built upon the Crystal GPU database query backend (Dense layout), and 4.3× with practical Sparse layout using less memory. Compared to current best compressed GPU bitmap implementation under perfect bin alignment (best case), our sparse layout achieves 1.4× speed end-to-end. We show that generic elementwise-style GPU fusion achieves only 1.34× speed, while our pipeline reaches 3.19× with 0.8% overhead versus dedicated compile-time kernels. Results on an NVIDIA H800 server GPU show the approach remains stable across GPU architectures, with smaller but still consistent fusion benefits on server GPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2e5df22-fe51-4ddc-8b39-a388e59d3d3bBuilds on12
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 112 citations
- Query Processing on Tensor Computation RuntimesDong He, Supun Chathuranga Nakandala, Dalitso Banda, Rathijit Sen et al.VLDB 2022 · 54 citations
- Tile-based Lightweight Integer Compression in GPUAnil Shanbhag, Bobbi W. Yogatama, Xiangyao Yu, Samuel MaddenSIGMOD 2022 · 45 citations
- GPU Database Systems Characterization and OptimizationJiashen Cao, Rathijit Sen, Matteo Interlandi, Joy Arulraj et al.VLDB 2024 · 36 citations
- MG-Join: A Scalable Join for Massively Parallel Multi-GPU ArchitecturesJohns Paul, Shengliang Lu, Bingsheng He, Chiew Tong LauSIGMOD 2021 · 31 citations
Related papers
- RABIT: Efficient Range Queries with Bitmap IndexingJunchang Wang, Fu Xiao, Manos AthanassoulisSIGMOD 2026
- RTIndeX: Exploiting Hardware-Accelerated GPU Raytracing for Database IndexingJustus Henneberg, Felix SchuhknechtVLDB 2023 · 25 citations
- More Bang for Your Buck(et): Fast and Space-Efficient Hardware-Accelerated Coarse-Granular Indexing on GPUsJustus Henneberg, Felix Martin Schuhknecht, Rosina Kharal, Trevor BrownICDE 2025 · 3 citations
- A Case for Graphics-driven Query ProcessingHarish Doraiswamy, Vikas Kalagi, Karthik Ramachandra, Jayant R. HaritsaVLDB 2023 · 4 citations
- CUBIT: Concurrent Updatable Bitmap IndexingJunchang Wang, Manos AthanassoulisVLDB 2025 · 7 citations
