F-TADOC: FPGA-Based Text Analytics Directly on Compression with HLS
Yanliang Zhou, Feng Zhang, Tuo Lin, Yuanjie Huang, Saiqin Long, Jidong Zhai, Xiaoyong Du
Abstract
With the development of loT and edge computing, data analytics on edge has become popular, and text analytics directly on compression (TADOC) has been proven to be a promising technology for edge data analytics. At the same time, Field Programmable Gate Array (FPGA) also has broad application prospects in data analytics systems. Unfortunately, there is no work to date showing how to support TADOC using FPGAs. We propose FPGA-based text analytics directly on compression with HLS, namely F - TADOC, which is the first framework using HLS to provide FPGA-based text analytics directly on compressed data. It effectively supports efficient text analytics on FPGA without decompressing input data. F-TADOC addresses three major challenges. First, TADOC involves a large number of dependencies with unbalanced workload of rules, which causes extremely low pipeline efficiency on FPG As. To solve it, we use layer-wise approach to traverse the DAG composed of rules and allocate different pipeline processing strategies for rules of different sizes. Second, the data volume required can be large that beyond the on-chip memory capacity of FPGAs. We develop a memory pool supporting hash structure and on-chip caches on FPGA to deal with this challenge. Third, when traversing the DAG, there are massive indirect addressing with a large number of random accesses. This leads to redundant time overhead caused by the latency in accessing the High Bandwidth Memory (HBM) during the pipeline. We optimize the F - TADOC algorithm by using dataflow to expand the nested loop, thus eliminate indirect addressing. With four widely used datasets, experiments show that F - TADOC achieves 4.63 x and 1.49 x performance speedup over TADOC and G- TADOC.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b6152cdf-45ba-46c3-a7d2-287135329c66Related papers
- G-TADOC: Enabling Efficient GPU-Based Text Analytics without DecompressionFeng Zhang, Zaifeng Pan, Yanliang Zhou, Jidong Zhai et al.ICDE 2021 · 31 citations
- Enabling Efficient NVM-Based Text Analytics without DecompressionXiaokun Fang, Feng Zhang, Junxiang Nong, Mingxing Zhang et al.ICDE 2024 · 1 citation
- : Near-Storage Accelerator for High-Performance Log AnalyticsSeongyoung Kang, Jiyoung An, Jinpyo Kim, Sang-Woo JunMICRO 2021 · 9 citations
- Homomorphic Compression: Making Text Processing on Compression UnlimitedJiawei Guan, Feng Zhang, Siqi Ma, Kuangyu Chen et al.SIGMOD 2024 · 11 citations
- Tile-based Lightweight Integer Compression in GPUAnil Shanbhag, Bobbi W. Yogatama, Xiangyao Yu, Samuel MaddenSIGMOD 2022 · 45 citations
