Fleet: A Framework for Massively Parallel Streaming on FPGAs
James Thomas, Pat Hanrahan, Matei Zaharia
Abstract
We present Fleet, a framework that offers a massively parallel streaming model for FPGAs and is effective in a number of domains well-suited for FPGA acceleration, including parsing, compression, and machine learning. Fleet requires the user to specify RTL for a processing unit that serially processes every input token in a stream, a far simpler task than writing a parallel processing unit. It then takes the user's processing unit and generates a hardware design with many copies of the unit as well as memory controllers to feed the units with separate streams and drain their outputs. Fleet includes a Chisel-based processing unit language. The language maintains Chisel's low-level performance control while adding a few productivity features, including automatic handling of ready-valid signaling and a native and automatically pipelined BRAM type. We evaluate Fleet on six different applications, including JSON parsing and integer compression, fitting hundreds of Fleet processing units on the Amazon F1 FPGA and outperforming CPU implementations by over 400× and GPU implementations by over 9× in performance per watt while requiring a similar number of lines of code.
• Hardware → Reconfigurable logic applications; Hardware description languages and compilation; Application specific processors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 20c5cf15-1da1-497c-8e3f-7df3476cfb20Cited by top-tier papers6
- Allo: A Programming Model for Composable Accelerator DesignHongzheng Chen, Niansong Zhang, Shaojie Xiang, Zhichen Zeng et al.PLDI 2024 · 41 citations
- PLD: fast FPGA compilation to make reconfigurable acceleration compatible with modern incremental refinement software developmentYuanlong Xiao, Eric Micallef, Andrew Butt, Matthew Hofmann et al.ASPLOS 2022 · 17 citations
- Skew-Oblivious Data Routing for Data Intensive Applications on FPGAs with HLSXinyu Chen, Hongshi Tan, Yao Chen, Bingsheng He et al.DAC 2021 · 6 citations
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMsHanchen Ye, Deming ChenMICRO 2025 · 5 citations
- Revet: A Language and Compiler for Dataflow ThreadsAlexander C. Rucker, Shiv Sundram, Coleman Smith, Matthew Vilim et al.HPCA 2024 · 3 citations
Related papers
- ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw DataElias Stehle, Hans-Arno JacobsenVLDB 2020 · 11 citations
- Type-directed scheduling of streaming acceleratorsDavid Durst, Matthew Feldman, Dillon Huff, David Akeley et al.PLDI 2020 · 49 citations
- Chronos: Efficient Speculative Parallelism for AcceleratorsMaleen Abeydeera, Daniel SánchezASPLOS 2020 · 32 citations
- Enabling Transparent Acceleration of Big Data Frameworks using Heterogeneous HardwareMaria Xekalaki, Juan Fumero, Athanasios Stratikopoulos, Katerina Doka et al.VLDB 2022 · 12 citations
- SPAGHETTI: Streaming Accelerators for Highly Sparse GEMM on FPGAsReza Hojabr, Ali Sedaghati, Amirali Sharifian, Ahmad Khonsari et al.HPCA 2021 · 66 citations
