SARIS: Accelerating Stencil Computations on Energy-Efficient RISC-V Compute Clusters with Indirect Stream Registers
Paul Scheffler, Luca Colagrande, Luca Benini
Abstract
Stencil codes are performance-critical in many compute-intensive applications, but suffer from significant address calculation and irregular memory access overheads. This work presents SARIS, a general and highly flexible methodology for stencil acceleration using register-mapped indirect streams. We demonstrate SARIS for various stencil codes on an eight-core RISC-V compute cluster with indirect stream registers, achieving significant speedups of 2.72x, near-ideal FPU utilizations of 81%, and energy efficiency improvements of 1.58x over an RV32G baseline on average. Scaling out to a 256-core manycore system, we estimate an average FPU utilization of 64%, an average speedup of 2.14x, and up to 15% higher fractions of peak compute than a leading GPU code generator.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Fast stencil-code computation on a wafer-scale processorKamil Rocki, Dirk Van Essendelft, Ilya Sharapov, Robert Schreiber et al.SC 2020 · 69 citations
- Unlimited Vector Extension with Data Streaming SupportJoao Mario Domingos, Nuno Neves, Nuno Roma, Pedro TomásISCA 2021 · 31 citations
- Scalable Distributed High-Order Stencil ComputationsMathias Jacquelin, Mauricio Araya-Polo, Jie MengSC 2022 · 17 citations
Related papers
- DiAG: a dataflow-inspired architecture for general-purpose processorsDong Kai Wang, Nam Sung KimASPLOS 2021 · 8 citations
- SARA: Scaling a Reconfigurable Dataflow AcceleratorYaqi Zhang, Nathan Zhang, Tian Zhao, Matt Vilim et al.ISCA 2021 · 58 citations
- R2D2: Removing ReDunDancy Utilizing Linearity of Address Generation in GPUsDongho Ha, Yunho Oh, Won Woo RoISCA 2023 · 9 citations
- RASA: Efficient Register-Aware Systolic Array Matrix Engine for CPUGeonhwa Jeong, Eric Qin, Ananda Samajdar, Christopher J. Hughes et al.DAC 2021 · 21 citations
- WASP: Exploiting GPU Pipeline Parallelism with Hardware-Accelerated Automatic Warp SpecializationNeal Clayton Crago, Sana Damani, Karthikeyan Sankaralingam, Stephen W. KecklerHPCA 2024 · 13 citations
