G-TADOC: Enabling Efficient GPU-Based Text Analytics without Decompression
Feng Zhang, Zaifeng Pan, Yanliang Zhou, Jidong Zhai, Xipeng Shen, Onur Mutlu, Xiaoyong Du
摘要
Text analytics directly on compression (TADOC) has proven to be a promising technology for big data analytics. GPUs are extremely popular accelerators for data analytics systems. Unfortunately, no work so far shows how to utilize GPUs to accelerate TADOC. We describe G-TADOC, the first framework that provides GPU-based text analytics directly on compression, effectively enabling efficient text analytics on GPUs without decompressing the input data.
G-TADOC solves three major challenges. First, TADOC involves a large amount of dependencies, which makes it difficult to exploit massive parallelism on a GPU. We develop a novel fine-grained thread-level workload scheduling strategy for GPU threads, which partitions heavily-dependent loads adaptively in a fine-grained manner. Second, in developing G-TADOC, thousands of GPU threads writing to the same result buffer leads to inconsistency while directly using locks and atomic operations lead to large synchronization overheads. We develop a memory pool with thread-safe data structures on GPUs to handle such difficulties. Third, maintaining the sequence information among words is essential for lossless compression. We design a sequencesupport strategy, which maintains high GPU parallelism while ensuring sequence information.
Our experimental evaluations show that G-TADOC provides 31.1× average speedup compared to state-of-the-art TADOC.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Generalizing Reuse Patterns for Efficient DNN on MicrocontrollersJiesong Liu, Bin Ren, Xipeng ShenASPLOS 2025 · 被引用 2 次
- Enabling Efficient NVM-Based Text Analytics without DecompressionXiaokun Fang, Feng Zhang, Junxiang Nong, Mingxing Zhang 等ICDE 2024 · 被引用 1 次
- Faith: An Efficient Framework for Transformer Verification on GPUsBoyuan Feng, Tianqi Tang, Yuke Wang, Zhaodong Chen 等USENIX ATC 2022
- Enabling Homomorphic Analytical Operations on Compressed Scientific Data with Multi-Stage DecompressionXuan Wu, Sheng Di, Tripti Agarwal, Kai Zhao 等ICDE 2026
它引用的顶会 Paper2
- FineStream: Fine-Grained Window-Based Stream Processing on CPU-GPU Integrated ArchitecturesFeng Zhang, Lin Yang, Shuhao Zhang, Bingsheng He 等USENIX ATC 2020 · 被引用 41 次
- Enabling Efficient Random Access to Hierarchically-Compressed DataFeng Zhang, Jidong Zhai, Xipeng Shen, Onur Mutlu 等ICDE 2020 · 被引用 22 次
相关 Paper
- F-TADOC: FPGA-Based Text Analytics Directly on Compression with HLSYanliang Zhou, Feng Zhang, Tuo Lin, Yuanjie Huang 等ICDE 2024 · 被引用 1 次
- Optimizing Random Access to Hierarchically-Compressed Data on GPUFeng Zhang, Yihua Hu, Haipeng Ding, Zhiming Yao 等SC 2022 · 被引用 5 次
- Tile-based Lightweight Integer Compression in GPUAnil Shanbhag, Bobbi W. Yogatama, Xiangyao Yu, Samuel MaddenSIGMOD 2022 · 被引用 45 次
- TempGraph: An Efficient Chain-driven Temporal Graph Computing Framework on the GPUJin Zhao, Qian Wang, Ligang He, Yu Zhang 等ASPLOS 2025
- ShadowVM: accelerating data plane for data analytics with bare metal CPUs and GPUsZhifang Li, Mingcong Han, Shangwei Wu, Chuliang WengPPoPP 2021 · 被引用 2 次
