Fleet: A Framework for Massively Parallel Streaming on FPGAs
James Thomas, Pat Hanrahan, Matei Zaharia
摘要
We present Fleet, a framework that offers a massively parallel streaming model for FPGAs and is effective in a number of domains well-suited for FPGA acceleration, including parsing, compression, and machine learning. Fleet requires the user to specify RTL for a processing unit that serially processes every input token in a stream, a far simpler task than writing a parallel processing unit. It then takes the user's processing unit and generates a hardware design with many copies of the unit as well as memory controllers to feed the units with separate streams and drain their outputs. Fleet includes a Chisel-based processing unit language. The language maintains Chisel's low-level performance control while adding a few productivity features, including automatic handling of ready-valid signaling and a native and automatically pipelined BRAM type. We evaluate Fleet on six different applications, including JSON parsing and integer compression, fitting hundreds of Fleet processing units on the Amazon F1 FPGA and outperforming CPU implementations by over 400× and GPU implementations by over 9× in performance per watt while requiring a similar number of lines of code.
• Hardware → Reconfigurable logic applications; Hardware description languages and compilation; Application specific processors.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Allo: A Programming Model for Composable Accelerator DesignHongzheng Chen, Niansong Zhang, Shaojie Xiang, Zhichen Zeng 等PLDI 2024 · 被引用 41 次
- PLD: fast FPGA compilation to make reconfigurable acceleration compatible with modern incremental refinement software developmentYuanlong Xiao, Eric Micallef, Andrew Butt, Matthew Hofmann 等ASPLOS 2022 · 被引用 17 次
- Skew-Oblivious Data Routing for Data Intensive Applications on FPGAs with HLSXinyu Chen, Hongshi Tan, Yao Chen, Bingsheng He 等DAC 2021 · 被引用 6 次
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMsHanchen Ye, Deming ChenMICRO 2025 · 被引用 5 次
- Revet: A Language and Compiler for Dataflow ThreadsAlexander C. Rucker, Shiv Sundram, Coleman Smith, Matthew Vilim 等HPCA 2024 · 被引用 3 次
相关 Paper
- ParPaRaw: Massively Parallel Parsing of Delimiter-Separated Raw DataElias Stehle, Hans-Arno JacobsenVLDB 2020 · 被引用 11 次
- Type-directed scheduling of streaming acceleratorsDavid Durst, Matthew Feldman, Dillon Huff, David Akeley 等PLDI 2020 · 被引用 49 次
- Chronos: Efficient Speculative Parallelism for AcceleratorsMaleen Abeydeera, Daniel SánchezASPLOS 2020 · 被引用 32 次
- Enabling Transparent Acceleration of Big Data Frameworks using Heterogeneous HardwareMaria Xekalaki, Juan Fumero, Athanasios Stratikopoulos, Katerina Doka 等VLDB 2022 · 被引用 12 次
- SPAGHETTI: Streaming Accelerators for Highly Sparse GEMM on FPGAsReza Hojabr, Ali Sedaghati, Amirali Sharifian, Ahmad Khonsari 等HPCA 2021 · 被引用 66 次
