Improving GPU Tensor Query Processing for Resource-Constrained Environments
Qian Xu, Feng Zhang, Shijie Gao, Kun Chen, Jianhua Wang, Zheng Chen, Xiaoyong Du
Abstract
Over the past decade, artificial intelligence (AI) has made remarkable strides, driven by researchers, engineers, and vendors who have relentlessly developed high-performance libraries across heterogeneous hardware and software. Recently, researchers in the database (DB) community have begun to harness this progress by representing data as tensors and redefining large-scale data operations through tensor-based computations. This approach seeks to leverage the rapid advancements in AI technologies. Typically, AI computing devices are heterogeneous, with limited and fixed on-device memory. We identify the lack of ability to process large-scale data as the biggest challenge in bridging the gap between the worlds of AI and databases. We propose TensorSlim, a memory-efficient big data analytical system incorporating three key innovations. First, we build a highly optimized operator execution engine that keeps GPU usage bounded by reusing buffers and releasing intermediates early. Second, we propose a tensor-driven processing framework that treats peak GPU memory as a first-class constraint and performs operator-aware GPU memory budgeting to guide query partitioning. Third, we develop a dependency-aware hybrid CPU–GPU query optimizer to manage dataflow dependencies and device placements, avoiding large GPU-resident intermediates. We demonstrate that TensorSlim can deliver 25.5× faster query speeds than traditional GPU-based solutions while achieving 6.4× speedup over CPU-based solutions on average. Additionally, TensorSlim reduces peak memory usage by 73.5%.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 65000541-ea93-417d-936f-c268ad7fb9f0Related papers
- Query Processing on Tensor Computation RuntimesDong He, Supun Chathuranga Nakandala, Dalitso Banda, Rathijit Sen et al.VLDB 2022 · 54 citations
- TQEx: Tensor-based Query Engine Enhanced by Bridging the GapHaitao Zhang, Ran Pang, Yuanyuan Zhu, Hao Zhang et al.SIGMOD 2026
- Scaling GPU-Accelerated Databases beyond GPU Memory SizeYinan Li, Bailu Ding, Ziyun Wei, Lukas M. Maas et al.VLDB 2025 · 7 citations
- TCUDB: Accelerating Database with Tensor ProcessorsYu-Ching Hu, Yuliang Li, Hung-Wei TsengSIGMOD 2022 · 32 citations
- Orchestrating Data Placement and Query Execution in Heterogeneous CPU-GPU DBMSBobbi W. Yogatama, Weiwei Gong, Xiangyao YuVLDB 2022 · 45 citations
