ngAP: Non-blocking Large-scale Automata Processing on GPUs
Tianao Ge, Tong Zhang, Hongyuan Liu
Abstract
Finite automata serve as compute kernels for various applications that require high throughput. However, despite the increasing compute power of GPUs, their potential in processing automata remains underutilized. In this work, we identify three major challenges that limit GPU throughput. 1) The available parallelism is insufficient, resulting in underutilized GPU threads. 2) Automata workloads involve significant redundant computations since a portion of states matches with repeated symbols. 3) The mapping between threads and states is switched dynamically, leading to poor data locality. Our key insight is that processing automata "one-symbol-at-a-time" serializes the execution, and thus needs to be revamped. To address these challenges, we propose Non-blocking Automata Processing, which allows parallel processing of different symbols in the input stream and also enables further optimizations: 1) We prefetch a portion of computations to increase the chances of processing multiple symbols simultaneously, thereby utilizing GPU threads better. 2) To reduce redundant computations, we store repeated computations in a memoization table, enabling us to substitute them with table lookups. 3) We privatize some computations to preserve the mapping between threads and states, thus improving data locality. Experimental results demonstrate that our approach outperforms the state-of-the-art GPU automata processing engine by an average of 7.9× and up to 901× across 20 applications.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get aa950faf-6cd3-41f4-b438-035e5de429e8Cited by top-tier papers4
- HybridSA: GPU Acceleration of Multi-pattern Regex Matching using Bit ParallelismAlexis Le Glaunec, Lingkun Kong, Konstantinos MamourasOOPSLA 2024 · 6 citations
- Interleaved Bitstream Execution for Multi-Pattern Regex Matching on GPUsTianao Ge, Xiaowen Chu, Hongyuan LiuMICRO 2025 · 4 citations
- Static Analysis for Efficient Streaming TokenizationAngela W. Li, Yudi Yang, Konstantinos MamourasASPLOS 2026 · 1 citation
- TempGraph: An Efficient Chain-driven Temporal Graph Computing Framework on the GPUJin Zhao, Qian Wang, Ligang He, Yu Zhang et al.ASPLOS 2025
Related papers
- Why GPUs are Slow at Executing NFAs and How to Make them FasterHongyuan Liu, Sreepathi Pai, Adwait JogASPLOS 2020 · 31 citations
- Hopps: Leveraging Sparsity to Accelerate Automata ProcessingXingran Du, Joel S. Emer, Daniel SánchezASPLOS 2025 · 1 citation
- Impala: Algorithm/Architecture Co-Design for In-Memory Multi-Stride Pattern MatchingElaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea Stan et al.HPCA 2020 · 43 citations
- Sunder: Enabling Low-Overhead and Scalable Near-Data Pattern Matching AccelerationElaheh Sadredini, Reza Rahimi, Mohsen Imani, Kevin SkadronMICRO 2021 · 11 citations
- FlexAmata: A Universal and Efficient Adaption of Applications to Spatial Automata Processing AcceleratorsElaheh Sadredini, Reza Rahimi, Marzieh Lenjani, Mircea Stan et al.ASPLOS 2020 · 23 citations
