TCUDB: Accelerating Database with Tensor Processors
Yu-Ching Hu, Yuliang Li, Hung-Wei Tseng
Abstract
The emergence of novel hardware accelerators has powered the tremendous growth of machine learning in recent years. These accelerators deliver incomparable performance gains in processing high-volume matrix operators, particularly matrix multiplication, a core component of neural network training and inference. In this work, we explored opportunities of accelerating database systems using NVIDIA's Tensor Core Units (TCUs). We present TCUDB, a TCU-accelerated query engine processing a set of query operators including natural joins and group-by aggregates as matrix operators within TCUs. Matrix multiplication was considered inefficient in the past; however, this strategy has remained largely unexplored in conventional GPU-based databases, which primarily rely on vector or scalar processing. We demonstrate the significant performance gain of TCUDB in a range of real-world applications including entity matching, graph query processing, and matrix-based data analytics. TCUDB achieves up to 288x speedup compared to a baseline GPU-based query engine.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fda15c9-4a7f-46fc-9dbe-07c95cfb36ecCited by top-tier papers12
- Query Processing on Tensor Computation RuntimesDong He, Supun Chathuranga Nakandala, Dalitso Banda, Rathijit Sen et al.VLDB 2022 · 54 citations
- JoinBoost: Grow Trees Over Normalized Data Using Only SQLZezhou Huang, Rathijit Sen, Jiaxiang Liu, Eugene WuVLDB 2023 · 23 citations
- Efficiently Processing Joins and Grouped Aggregations on GPUsBowen Wu, Dimitrios Koutsoukos, Gustavo AlonsoSIGMOD 2025 · 15 citations
- BladeDISC: Optimizing Dynamic Shape Machine Learning Workloads via Compiler ApproachZhen Zheng, Zaifeng Pan, Dalin Wang, Kai Zhu et al.SIGMOD 2024 · 14 citations
- Vortex: Overcoming Memory Capacity Limitations in GPU-Accelerated Large-Scale Data AnalyticsYichao Yuan, Advait Iyer, Lin Ma, Nishil TalatiVLDB 2025 · 11 citations
Builds on4
- A Study of the Fundamental Performance Characteristics of GPUs and CPUs for Database AnalyticsAnil Shanbhag, Samuel Madden, Xiangyao YuSIGMOD 2020 · 112 citations
- Pump Up the Volume: Processing Large Data on GPUs with Fast InterconnectsClemens Lutz, Sebastian Breß, Steffen Zeuch, Tilmann Rabl et al.SIGMOD 2020 · 99 citations
- Lowering the Latency of Data Processing Pipelines Through FPGA based Hardware AccelerationMuhsen Owaida, Gustavo Alonso, Laura Fogliarini, Anthony Hock-Koon et al.VLDB 2020 · 48 citations
- A Relational Matrix Algebra and its Implementation in a Column StoreOksana Dolmatova, Nikolaus Augsten, Michael H. BöhlenSIGMOD 2020 · 11 citations
Related papers
- RayDB: Building Databases with Ray Tracing CoresXuri Shi, Kai Zhang, X. Sean Wang, Xiaodong Zhang et al.VLDB 2026 · 3 citations
- Acc-SpMM: Accelerating General-purpose Sparse Matrix-Matrix Multiplication with GPU Tensor CoresHaisha Zhao, San Li, Jiaheng Wang, Chunbao Zhou et al.PPoPP 2025 · 18 citations
- TQEx: Tensor-based Query Engine Enhanced by Bridging the GapHaitao Zhang, Ran Pang, Yuanyuan Zhu, Hao Zhang et al.SIGMOD 2026
- High Accuracy Matrix Computations on Neural Engines: A Study of QR Factorization and its ApplicationsShaoshuai Zhang, Elaheh Baharlouei, Panruo WuHPDC 2020 · 16 citations
- DTC-SpMM: Bridging the Gap in Accelerating General Sparse Matrix Multiplication with Tensor CoresRuibo Fan, Wei Wang, Xiaowen ChuASPLOS 2024 · 46 citations
