Lune

PPoPP2026Top-tier venue

SPIDER: Unleashing Sparse Tensor Cores for Stencil Computation via Strided Swapping

Qiqi Gu, Chenpeng Wu, Heng Shi, Jianguo Yao

2026Year
1Citations

Abstract

Recent research has focused on accelerating stencil computations by exploiting emerging hardware like Tensor Cores. To leverage these accelerators, the stencil operation must be transformed to matrix multiplications. However, this transformation introduces undesired sparsity into the kernel matrix, leading to significant redundant computation.

In this paper, we present SPIDER, the first system to turn this unresolved sparsity into an optimization opportunity by exploring the potential of Sparse Tensor Cores (SpTCs) for stencil acceleration. Specifically, SPIDER introduces an efficient and elegant transformation method that integrates two cooperative techniques: an ahead-of-time strided swapping transformation for kernel matrices and an on-the-fly rowswapping mechanism for inputs. This rule-based approach effectively transforms stencil computation into operations compatible with SpTCs, introducing only slight compile-time overhead and zero runtime overhead. Additionally, SPIDER incorporates multiple optimizations to maximize computational efficiency. Experimental evaluations demonstrate that SPIDER outperforms vendor library cuDNN by 6.20× and state-of-the-art (SOTA) Tensor Core-based approaches (Con-vStencil, FlashFFTStencil, etc.) by 2.00× on average.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext dbd818d7-12af-4761-90df-0cda57241ee8

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines