Exploiting Computation Reuse for Stencil Accelerators
Yuze Chi, Jason Cong
摘要
Stencil kernel is an important type of kernel used extensively in many application domains. Over the years, researchers have been studying the optimizations on parallelization, communication reuse, and computation reuse for various target platforms. However, challenges still exist, especially on the computation reuse problem for accelerators, due to the lack of complete design-space exploration and effective design-space pruning. In this paper, we present solutions to the above challenges for a wide range of stencil kernels (i.e., stencil with reduction operations), where the computation reuse patterns are extremely flexible due to the commutative and associative properties. We formally define the complete design space, based on which we present a provably optimal dynamic programming algorithm and a heuristic beam search algorithm that provides near-optimal solutions under an architecture-aware model. Experimental results show that for synthesizing stencil kernels to FPGAs, compared with state-of-the-art stencil compiler without computation reuse capability, our proposed algorithm can reduce the look-up table (LUT) and digital signal processor (DSP) usage by 58.1% and 54.6% on average respectively, which leads to an average speedup of 2.3× for compute-intensive kernels, outperforming the latest CPU/GPU results.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Analysis and Optimization of the Implicit Broadcasts in FPGA HLS to Improve Maximum FrequencyLicheng Guo, Jason Lau, Yuze Chi, Jie Wang 等DAC 2020 · 被引用 21 次
- Redundant Array Computation EliminationZixuan Wang, Liang Yuan, Xianmeng Jiang, Kun Li 等PLDI 2026
相关 Paper
- Moirae: Generating High-Performance Composite Stencil Programs with Global OptimizationsXiaoyan Liu, Xinyu Yang, Kejie Ma, Shanghao Liu 等SC 2024 · 被引用 2 次
- Reducing redundancy in data organization and arithmetic calculation for stencil computationsKun Li, Liang Yuan, Yunquan Zhang, Yue YueSC 2021 · 被引用 12 次
- ETO: Accelerating Optimization of DNN Operators by High-Performance Tensor Program ReuseJingzhi Fang, Yanyan Shen, Yue Wang, Lei ChenVLDB 2022 · 被引用 10 次
- A Streaming Collectives Interface Targeting Dataflow Acceleration and HPC WorkloadsNicholas Contini, Jake Queiser, Bharath Ramesh, Hari Subramoni 等SC 2025 · 被引用 1 次
- StreamTensor: Make Tensors Stream in Dataflow Accelerators for LLMsHanchen Ye, Deming ChenMICRO 2025 · 被引用 5 次
