SC2020Top-tier venue
fBLAS: streaming linear algebra on FPGA
Tiziano De Matteis, Johannes de Fine Licht, Torsten Hoefler
Abstract
Spatial computing architectures pose an attractive alternative to mitigate control and data movement overheads typical of load-store architectures. In practice, these devices are rarely considered in the HPC community due to the steep learning curve, low productivity, and the lack of available libraries for fundamental operations. High-level synthesis (HLS) tools are facilitating hardware programming, but optimizing for these architectures requires factoring in new transformations and resources/performance trade-offs. We present FBLAS, an opensource HLS implementation of BLAS for FPGAs, that enables reusability, portability and easy integration with existing software and hardware codes. FBLAS' implementation allows scaling hardware modules to exploit on-chip resources, and module interfaces are designed to natively support streaming on-chip communications, allowing them to be composed to reduce offchip communication. With FBLAS, we set a precedent for FPGA library design, and contribute to the toolbox of customizable hardware components necessary for HPC codes to start productively targeting reconfigurable platforms.
Index Terms-Spatial architectures, high level synthesis, hardware library
• FBLAS, the first portable and open source BLAS implementation on FPGA, realized entirely with state-of-the-art
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07c26eb6-aefe-4fdb-aa55-de2eaf4bfb78Cited by top-tier papers4
- FlexTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous SystemSize Zheng, Yun Liang, Shuo Wang, Renze Chen et al.ASPLOS 2020 · 171 citations
- High Performance, Low Power Matrix Multiply Design on ACAP: from Architecture, Design Challenges and DSE PerspectivesJinming Zhuang, Zhuoping Yang, Peipei ZhouDAC 2023 · 28 citations
- A submatrix-based method for approximate matrix function evaluation in the quantum chemistry code CP2KMichael Lass, Robert Schade, Thomas D. Kühne, Christian PlesslSC 2020 · 7 citations
- Misam: Machine Learning Assisted Dataflow Selection in Accelerators for Sparse Matrix MultiplicationSanjali Yadav, Amirmahdi Namjoo, Bahar AsgariMICRO 2025 · 6 citations
Related papers
- A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC WorkloadsNicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni et al.SC 2025 · 1 citation
- TensorLib: A Spatial Accelerator Generation Framework for Tensor AlgebraLiancheng Jia, Zizhang Luo, Liqiang Lu, Yun LiangDAC 2021 · 49 citations
- HiSpTRSV: Exploring Tile-Level Parallelism for SpTRSV Acceleration on FPGAsFan Sun, Fang Dong, Dian ShenDAC 2025
- OverGen: Improving FPGA Usability through Domain-specific Overlay GenerationSihao Liu, Jian Weng, Dylan Kupsh, Atefeh Sohrabizadeh et al.MICRO 2022 · 32 citations
- Predictable accelerator design with time-sensitive affine typesRachit Nigam, Sachille Atapattu, Samuel Thomas, Zhijing Li et al.PLDI 2020 · 58 citations
