F-TADOC: FPGA-Based Text Analytics Directly on Compression with HLS
Yanliang Zhou, Feng Zhang, Tuo Lin, Yuanjie Huang, Saiqin Long, Jidong Zhai, Xiaoyong Du
摘要
With the development of loT and edge computing, data analytics on edge has become popular, and text analytics directly on compression (TADOC) has been proven to be a promising technology for edge data analytics. At the same time, Field Programmable Gate Array (FPGA) also has broad application prospects in data analytics systems. Unfortunately, there is no work to date showing how to support TADOC using FPGAs. We propose FPGA-based text analytics directly on compression with HLS, namely F - TADOC, which is the first framework using HLS to provide FPGA-based text analytics directly on compressed data. It effectively supports efficient text analytics on FPGA without decompressing input data. F-TADOC addresses three major challenges. First, TADOC involves a large number of dependencies with unbalanced workload of rules, which causes extremely low pipeline efficiency on FPG As. To solve it, we use layer-wise approach to traverse the DAG composed of rules and allocate different pipeline processing strategies for rules of different sizes. Second, the data volume required can be large that beyond the on-chip memory capacity of FPGAs. We develop a memory pool supporting hash structure and on-chip caches on FPGA to deal with this challenge. Third, when traversing the DAG, there are massive indirect addressing with a large number of random accesses. This leads to redundant time overhead caused by the latency in accessing the High Bandwidth Memory (HBM) during the pipeline. We optimize the F - TADOC algorithm by using dataflow to expand the nested loop, thus eliminate indirect addressing. With four widely used datasets, experiments show that F - TADOC achieves 4.63 x and 1.49 x performance speedup over TADOC and G- TADOC.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- G-TADOC: Enabling Efficient GPU-Based Text Analytics without DecompressionFeng Zhang, Zaifeng Pan, Yanliang Zhou, Jidong Zhai 等ICDE 2021 · 被引用 31 次
- Enabling Efficient NVM-Based Text Analytics without DecompressionXiaokun Fang, Feng Zhang, Junxiang Nong, Mingxing Zhang 等ICDE 2024 · 被引用 1 次
- : Near-Storage Accelerator for High-Performance Log AnalyticsSeongyoung Kang, Jiyoung An, Jinpyo Kim, Sang-Woo JunMICRO 2021 · 被引用 9 次
- Homomorphic Compression: Making Text Processing on Compression UnlimitedJiawei Guan, Feng Zhang, Siqi Ma, Kuangyu Chen 等SIGMOD 2024 · 被引用 11 次
- Tile-based Lightweight Integer Compression in GPUAnil Shanbhag, Bobbi W. Yogatama, Xiangyao Yu, Samuel MaddenSIGMOD 2022 · 被引用 45 次
