Lune

ICDE2026顶会

Improving GPU Tensor Query Processing for Resource-Constrained Environments

Qian Xu, Feng Zhang, Shijie Gao, Kun Chen, Jianhua Wang, Zheng Chen, Xiaoyong Du

2026年份

摘要

Over the past decade, artificial intelligence (AI) has made remarkable strides, driven by researchers, engineers, and vendors who have relentlessly developed high-performance libraries across heterogeneous hardware and software. Recently, researchers in the database (DB) community have begun to harness this progress by representing data as tensors and redefining large-scale data operations through tensor-based computations. This approach seeks to leverage the rapid advancements in AI technologies. Typically, AI computing devices are heterogeneous, with limited and fixed on-device memory. We identify the lack of ability to process large-scale data as the biggest challenge in bridging the gap between the worlds of AI and databases. We propose TensorSlim, a memory-efficient big data analytical system incorporating three key innovations. First, we build a highly optimized operator execution engine that keeps GPU usage bounded by reusing buffers and releasing intermediates early. Second, we propose a tensor-driven processing framework that treats peak GPU memory as a first-class constraint and performs operator-aware GPU memory budgeting to guide query partitioning. Third, we develop a dependency-aware hybrid CPU–GPU query optimizer to manage dataflow dependencies and device placements, avoiding large GPU-resident intermediates. We demonstrate that TensorSlim can deliver 25.5× faster query speeds than traditional GPU-based solutions while achieving 6.4× speedup over CPU-based solutions on average. Additionally, TensorSlim reduces peak memory usage by 73.5%.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖