fBLAS: streaming linear algebra on FPGA
Tiziano De Matteis, Johannes de Fine Licht, Torsten Hoefler
摘要
Spatial computing architectures pose an attractive alternative to mitigate control and data movement overheads typical of load-store architectures. In practice, these devices are rarely considered in the HPC community due to the steep learning curve, low productivity, and the lack of available libraries for fundamental operations. High-level synthesis (HLS) tools are facilitating hardware programming, but optimizing for these architectures requires factoring in new transformations and resources/performance trade-offs. We present FBLAS, an opensource HLS implementation of BLAS for FPGAs, that enables reusability, portability and easy integration with existing software and hardware codes. FBLAS' implementation allows scaling hardware modules to exploit on-chip resources, and module interfaces are designed to natively support streaming on-chip communications, allowing them to be composed to reduce offchip communication. With FBLAS, we set a precedent for FPGA library design, and contribute to the toolbox of customizable hardware components necessary for HPC codes to start productively targeting reconfigurable platforms.
Index Terms-Spatial architectures, high level synthesis, hardware library
• FBLAS, the first portable and open source BLAS implementation on FPGA, realized entirely with state-of-the-art
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen 等ASPLOS 2020 · 被引用 171 次
- High Performance, Low Power Matrix Multiply Design on ACAP: from Architecture, Design Challenges and DSE PerspectivesJinming Zhuang, Zhuoping Yang, Peipei ZhouDAC 2023 · 被引用 28 次
- A submatrix-based method for approximate matrix function evaluation in the quantum chemistry code CP2KMichael Lass, Robert Schade, Thomas D. Kühne, Christian PlesslSC 2020 · 被引用 7 次
- Misam: Machine Learning Assisted Dataflow Selection in Accelerators for Sparse Matrix MultiplicationSanjali Yadav, Amirmahdi Namjoo, Bahar AsgariMICRO 2025 · 被引用 6 次
相关 Paper
- A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC WorkloadsNicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni 等SC 2025 · 被引用 1 次
- TensorLib: A Spatial Accelerator Generation Framework for Tensor AlgebraLiancheng Jia, Zizhang Luo, Liqiang Lu, Yun LiangDAC 2021 · 被引用 49 次
- HiSpTRSV: Exploring Tile-Level Parallelism for SpTRSV Acceleration on FPGAsFan Sun, Fang Dong, Dian ShenDAC 2025
- OverGen: Improving FPGA Usability through Domain-specific Overlay GenerationSihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh 等MICRO 2022 · 被引用 32 次
- Predictable accelerator design with time-sensitive affine typesRachit Nigam, Sachille Atapattu, Samuel Thomas, Zhijing Li 等PLDI 2020 · 被引用 58 次
