Exploiting Computation Reuse for Stencil Accelerators
Yuze Chi, Jason Cong
Abstract
Stencil kernel is an important type of kernel used extensively in many application domains. Over the years, researchers have been studying the optimizations on parallelization, communication reuse, and computation reuse for various target platforms. However, challenges still exist, especially on the computation reuse problem for accelerators, due to the lack of complete design-space exploration and effective design-space pruning. In this paper, we present solutions to the above challenges for a wide range of stencil kernels (i.e., stencil with reduction operations), where the computation reuse patterns are extremely flexible due to the commutative and associative properties. We formally define the complete design space, based on which we present a provably optimal dynamic programming algorithm and a heuristic beam search algorithm that provides near-optimal solutions under an architecture-aware model. Experimental results show that for synthesizing stencil kernels to FPGAs, compared with state-of-the-art stencil compiler without computation reuse capability, our proposed algorithm can reduce the look-up table (LUT) and digital signal processor (DSP) usage by 58.1% and 54.6% on average respectively, which leads to an average speedup of 2.3× for compute-intensive kernels, outperforming the latest CPU/GPU results.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get beaef626-874a-4d90-af9b-98c4df69e41eCited by top-tier papers2
- Analysis and Optimization of the Implicit Broadcasts in FPGA HLS to Improve Maximum FrequencyLicheng Guo, Jason Lau, Yuze Chi, Jie Wang et al.DAC 2020 · 21 citations
- Redundant Array Computation EliminationZixuan Wang, Liang Yuan, Xianmeng Jiang, Kun Li et al.PLDI 2026
Related papers
- Moirae: Generating High-Performance Composite Stencil Programs with Global OptimizationsXiaoyan Liu, Xinyu Yang, Kejie Ma, Shanghao Liu et al.SC 2024 · 2 citations
- Reducing redundancy in data organization and arithmetic calculation for stencil computationsKun Li, Liang Yuan, Yunquan Zhang, Yue YueSC 2021 · 12 citations
- ETO: Accelerating Optimization of DNN Operators by High-Performance Tensor Program ReuseJingzhi Fang, Yanyan Shen, Yue Wang, Lei ChenVLDB 2022 · 10 citations
- A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC WorkloadsNicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni et al.SC 2025 · 1 citation
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMsHanchen Ye, Deming ChenMICRO 2025 · 5 citations
